{"id":"43c4f188-f5ed-4f73-a99c-917800c08fa7","arxiv_id":"2608.08220","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors argue that RL agents should be classified as moral only relative to a distinct moral reward function, and use this criterion to assess three RL-based value alignment proposals.","lead":"This paper applies ideas from metanormative philosophy to reinforcement learning, arguing that calling an RL agent's behavior 'moral' requires a dedicated moral reward function as the standard. It then uses this framework to evaluate three existing RL-based approaches to machine ethics and value alignment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The dedicated-reward-function requirement is asserted, not derived; constraint-based encodings may ground the same normative vocabulary, so the paper's negative verdict on Noothigattu et al. is not forced.","rationale":"The strongest_claim is that moral vocabulary is legitimate only when a distinct moral reward function grounds a moral domain, and the paper uses this to reject constraint-based approaches such as Noothigattu et al. The reader's weakest_assumption correctly identifies that the requirement is asserted in Section 5 without argument. I agree. The claim is load-bearing because without it the paper's case-study verdicts lose their force and the promised 'clearer criteria' reduce to a modeling preference. The proposed test targets the precise point: can the same normative categories be defined relative to constraints rather than a reward function? In RL, constrained MDPs are a standard formalism; forbidden actions and required policies can be read as deontic facts. If that reading is coherent, then the paper's 'no principled basis' sentence fails, and the framework needs to be weakened to a sufficiency claim ('a distinct reward function is one way to ground a moral domain') rather than a necessity claim. The Abel et al. tension is a secondary symptom of the same unargued identification: the paper wants to separate 'maximizing expected reward' from 'maximizing moral reward', but with a single ethical utility function these coincide, so the criticism requires an external notion of motive that the framework did not introduce. Because the reader already flagged the main assumption and set CONDITIONAL, I do not move the verdict; I only sharpen the test.","tokens_in":16002,"tokens_out":4076,"duration_ms":39262,"concrete_test":"Formally model Noothigattu et al.'s Pac-Man domain as a constrained MDP without a separate moral reward function: keep the domain reward R, impose a hard constraint forbidding eating ghosts (e.g., action masking or constrained policy optimization), and define deontic categories via the constraint set—actions that violate the constraint are forbidden, compliant policies are permissible, and the constrained optimum is required. If the resulting system admits the same normative vocabulary from Section 5 and Table 2, then the paper's claim that there is 'no principled basis for the application of moral categories' in the absence of a distinct moral reward function is falsified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The framework's central criterion depends on the Section 5 assertion that representing a moral domain requires that 'the first thing one needs is a distinct reward function that can serve as a basis for the moral domain.' This is an engineering choice, not a consequence of metanormative theory. A normative domain in metanormative theory is fixed by a standard against which categories like right/wrong, good/bad, and fitting/unfitting apply; that standard need not be a separate scalar reward channel. Hard constraints, action prohibitions, and normative systems can supply such a standard. In a constrained MDP, deontic categories apply directly: actions violating the constraint are forbidden, compliant policies are permissible or required, and the constrained optimum is fitting. Thus Table 2's assignment of categories can be reproduced without a distinct reward function. The Noothigattu et al. critique in Section 6 therefore overreaches: it shows only that their blended rewards do not wear their moral content on their sleeve, not that no moral categories can be applied. There is also an internal tension in the Abel et al. discussion: if the 'true ethical utility function' is the POMDP reward, then maximizing expected reward is maximizing moral reward, so the claim that the learned policy is moral only 'incidentally' is unexplained. Both weaknesses trace to the unargued identification of moral domains with dedicated reward functions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a metanormative-theoretic framework for classifying and evaluating RL-based moral agents. It introduces normative categories from metanormative theory (deontic, evaluative, fittingness, and reason-based), distinguishes normative domains, and maps these categories onto components of RL systems (state-action pairs, policies, and reward signals). The central thesis is that moral vocabulary in RL contexts is legitimate only when a distinct moral reward function grounds a moral domain. The authors apply the framework to three existing approaches: Abel et al.'s POMDP-based ethical decision making, Noothigattu et al.'s policy orchestration with learned constraints, and Rodriguez-Soto et al.'s multi-objective RL. They argue that Abel et al. and Rodriguez-Soto et al. provide bases for moral categories (with weaknesses), while Noothigattu et al. do not.","tokens_in":16243,"tokens_out":5550,"duration_ms":47501,"significance":"If the framework were accepted, it would give AI researchers a principled vocabulary for distinguishing moral from merely goal-directed RL systems and a basis for comparing approaches. The paper's strength is its clear synthesis of a substantial philosophical literature and its concrete mapping in Table 2, which is a useful conceptual tool. The authors are also explicit about the simplified and contested nature of the philosophy. However, the central criterion—the requirement of a dedicated moral reward function—is stipulated rather than derived, and the negative verdicts on Abel et al. and Noothigattu et al. depend on this stipulation. The paper is therefore a promising starting point rather than a settled framework.","major_comments":[{"comment":"Section 5, paragraph beginning 'How should one represent the moral domain in an RL system?': The claim that 'the first thing one needs is a distinct reward function that can serve as a basis for the moral domain' is presented as 'the most natural and conservative answer' but is not derived from the metanormative theory presented in Sections 3–4. In that theory, a normative domain is fixed by a standard against which deontic, evaluative, and fittingness categories apply; nothing in the account requires that standard to be encoded as a scalar reward function. A constrained MDP with hard prohibitions supplies a standard: violated actions are forbidden, compliant policies are permissible, and the constrained optimum is fitting. The paper therefore needs either to argue that only reward functions can serve as such standards, or to restrict the criterion to RL systems that use reward-based normative encodings. As written, the criterion is an engineering choice, and the later negative verdict on Noothigattu et al. in Section 6 inherits this presupposition.","section":"Section 5"},{"comment":"The claim that the learned policy is moral only 'incidentally' appears to contradict the paper's own Section 5 mapping. If the POMDP's 'true ethical utility function' is the reward function that grounds the moral domain, then the optimal policy is, by the Section 5 definition, the morally best/right/fitting policy, and maximizing expected reward with respect to that function just is maximizing moral reward. The sentence 'the question of whether a policy is morally good, right, or fitting has little to do with maximizing expected reward signal (as opposed to maximizing moral reward signal)' conflates the two, because in this example the expected reward is the expected moral reward. The authors should clarify whether they take the 'true ethical utility function' to be the reward function actually used in optimization, or whether they intend a distinction between observable reward and underlying utility; the current text does not support the 'incidental' criticism.","section":"Section 6, Abel et al."},{"comment":"The statement that 'there is no principled basis for the application of moral categories in this context' overreaches. The learned constraint rewards in the πC policy define a normative standard: eating ghosts is assigned a strongly negative value, which can ground the deontic judgment that this action is forbidden within the constrained domain, and the constrained optimum can be called fitting. The problem the authors identify is better described as a blending of two domains (domain reward and moral penalty) that obscures which standard is in force, not as the absence of any basis for moral categories. This reformulation would preserve the criticism without making it depend on the unargued reward-function requirement.","section":"Section 6, Noothigattu et al."},{"comment":"The paper states that the moral status of an RL agent's behavior 'must ultimately depend on the moral status of some of the following: states; actions; reward signals; learning algorithms; and policies,' but the subsequent analysis in Section 5 and Table 2 only covers state-action pairs, policies, and reward signals; learning algorithms are never assigned normative categories. Since the list is presented as exhaustive of the possible loci of moral status, the omission leaves the framework incomplete with respect to its own central claim. The authors should either extend the analysis to learning algorithms or remove them from the list.","section":"Section 2 and Table 2"}],"minor_comments":[{"comment":"The word 'utalitarianism' should be 'utilitarianism'; similarly, 'POMPD' in Section 6 should be 'POMDP'.","section":"Section 2"},{"comment":"The sentence 'If normatively sensitive situations ar thought of' should read 'are thought of'.","section":"Section 3"},{"comment":"The sentence 'What is unusual.' is incomplete and appears to be a leftover from an intended contrast; the authors should complete or delete it.","section":"Section 6"},{"comment":"The table captions read 'T able' instead of 'Table'; please fix the formatting.","section":"Tables 1 and 3"},{"comment":"The entry for Abel et al. lists 'the agent's sole goal is to act morally' as a weakness, but the surrounding text criticizes the agent for learning only one of the morally optimal policies; spelling out why having only this goal is a weakness would make the table more informative.","section":"Section 6, Table 3"},{"comment":"The concluding caveat that the criterion 'should not be decisive' sits uneasily with the categorical claims in Section 6, such as 'there is no principled basis for the application of moral categories'; the authors should reconcile the hedged conclusion with the body's categorical verdicts.","section":"Section 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is a conceptual contribution and should be assessed on the cogency of its mapping from metanormative theory to RL, not on empirical validation. The main risk to the paper's claims is the unargued identification of moral domains with dedicated reward functions; the authors should be pushed to defend or qualify this identification. The paper fits well in an AI-ethics or value-alignment venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper imports metanormative theory into RL-based machine ethics and does it cleanly. The taxonomy of deontic, evaluative, fittingness, and reason-based categories, mapped onto state-action pairs, policies, and reward signals, is genuinely useful and I haven't seen it before in this form. The three case studies are fair and instructive, and the authors are honest about the simplifications they make and about disagreement within metanormative theory itself.\n\nThe evaluation of Rodriguez-Soto et al. is particularly balanced, and the discussion of how multiple reward functions can partially address the problem of moral reasons in a single scalar signal is a thoughtful point. The paper also does a good service in pressing the field to be more precise about what 'moral' and 'value' mean when attached to RL agents.\n\nThe main weakness is the load-bearing premise in Section 5: that a moral domain in an RL system requires a distinct reward function. The paper says this is the 'most natural and conservative' way to proceed, but it never argues that it is the only way. That matters because Section 6's verdict on Noothigattu et al. depends on treating that premise as necessary: their constrained policies are said to have 'no principled basis' for moral categories because the rewards blend domain and moral penalties. But a constrained MDP can supply a normative standard just as well: the constraint forbids certain actions, compliant policies are permissible or required, and deontic categories apply directly. The stress-test note is right that the negative verdict overreaches.\n\nThere is also an internal tension in the Abel et al. discussion. If the 'true ethical utility function' is the POMDP reward, then maximizing expected reward is maximizing moral utility, so calling the learned policy moral only 'incidental' seems inconsistent with the paper's own criterion.\n\nThese are real soft spots, but not fatal to the paper's value. The framework works well as a sufficient condition for when moral vocabulary is meaningful, and as a common language for comparing approaches. What it does not do is establish a necessary condition, and the paper sometimes slides between those two readings.\n\nThis paper deserves a serious referee. It is carefully written, philosophically literate, and opens a useful conversation. I would engage with it and bring it to a reading group to debate the reward-function assumption. Send it to peer review, but expect the necessity claim to be challenged and ask the authors to either defend it or scale it back.","headline":"A clear, useful mapping of metanormative categories onto RL components, but the core criterion that a moral domain requires a dedicated reward function is asserted rather than argued, and it drives a negative verdict on constrained-RL approaches that overreaches.","tokens_in":16747,"tokens_out":2830,"would_cite":true,"duration_ms":27702,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An RL agent's behavior counts as moral only when a distinct reward function grounds a moral domain.","keywords":["machine ethics","metanormative theory","reinforcement learning","value alignment","moral reward function","normative categories","RLHF","multi-objective RL"],"falsifier":"A controlled experiment in which an RL agent trained solely with human feedback, without any separate moral reward function, outperforms a moral-reward-trained agent on a battery of moral scenarios in new environments would directly contradict the claim that a distinct moral reward function is necessary for moral classification.","tokens_in":15787,"feed_emoji":"⚖️","tokens_out":10417,"duration_ms":83320,"temperature":0.7,"pith_summary":"This paper brings tools from metanormative theory—the branch of philosophy that studies the structure of normative domains—to bear on reinforcement-learning (RL) agents. It argues that an RL agent's behavior can be classified as moral only when a distinct reward function establishes a moral domain within the architecture, so that states, actions, and policies can be assessed as morally right, wrong, good, or bad against that standard. The paper derives this criterion by mapping the standard components of an RL system onto four families of normative categories: deontic, evaluative, fittingness, and reason-based. It then applies the criterion to three prominent RL ethics proposals, finding that one makes moral talk meaningful but treats moral behavior as incidental, another leaves no principled basis for moral classification, and a third is structurally well-suited but misplaces the prudential value of achievement and imposes a context-invariant ordering of moral reasons. If the framework is right, it gives researchers a principled way to say when a trained RL policy is genuinely moral rather than merely harm-avoiding.","feed_headline":"No moral reward function, no moral RL agent","feed_subtitle":"A framework from metanormative philosophy says which RL systems can be judged moral, and which cannot.","key_machinery":"The machinery that carries the argument is the pairing of a dedicated reward function with the taxonomy of normative categories from metanormative theory. The reward function—conceived as the 'basis for the moral domain'—supplies the standard against which normative categories apply, while the categories (deontic, evaluative, fittingness, and reason-based) provide the vocabulary used to qualify the components of an RL system: state-action pairs, policies, and reward signals. This mapping, summarized in the paper's Table 2, does the work of separating the moral domain from the prudential domain and yields the paper's evaluative verdicts on concrete RL proposals.","core_discovery":"The paper's central claim is that moral vocabulary applied to RL agents is legitimate only when the architecture contains a distinct moral reward function that grounds a moral domain separate from the domain of self-interest or prudence. Given such a function, policies can be morally best, right, and fitting, and state-action pairs can be morally good or bad; without it, there is no principled basis for applying moral categories to the system's behavior. The paper therefore reads the moral status of an RL agent's behavior as determined by the moral status of its states, actions, reward signals, learning algorithms, and policies, and it treats the deontic, evaluative, fittingness, and reason-based categories from metanormative theory as the vocabulary that makes that determination precise. On this view, a policy is moral when it maximizes the moral reward signal, and talk of 'moral behavior' is derivative from the morally optimal policy, just as 'rational behavior' is derivative from the reward-maximizing policy in the prudential domain.","pith_inferences":["The paper leaves implicit a testable design principle: two behaviorally identical policies could differ in moral status if one maximizes a moral reward and the other merely happens to avoid harm while maximizing a domain reward.","The paper's critique of scalar rewards as 'mere numbers' suggests that moral RL may need richer reward representations, such as vector-valued rewards with a lexicographic ordering, to capture different kinds of moral reasons.","The framework could be pressed on safety-layer and action-masking approaches, where moral constraints are implemented without any reward channel; the paper does not say whether such a filter would count as a moral domain, but its criterion suggests it would not.","The criterion could be operationalized as an audit: inspect the reward function of a deployed RL system and check whether a distinct moral reward channel exists before attributing moral status to its behavior."],"forward_implications":["If the framework is correct, RLHF-based agents that learn only from human preference labels, with no separate moral reward function, cannot be classified as moral under a principled criterion.","Constraint-based approaches that blend domain rewards with ethical penalties produce policies that belong to neither the moral nor the rational domain, so calling them 'ethical' is ungrounded.","Multi-objective RL with a vector of ethical reward functions is structurally compatible with moral categories, but only if the value of achievement is placed in the prudential domain and the relative weight of moral reasons is allowed to vary across contexts.","An RL policy learned by maximizing expected reward is not thereby moral; the reward function being optimized must be the moral reward function for moral vocabulary to apply.","Value talk in AI becomes precise only when translated into normative categories; otherwise claims about 'values' in alignment research lack a clear normative basis."],"supporting_citations":[{"why":"Supplies the taxonomy of deontic, evaluative, and fittingness categories and the features that distinguish them.","marker":"[9]"},{"why":"Provides the overall versus contributory distinction and the treatment of normative domains and the weight of reasons.","marker":"[27]"},{"why":"Defines fittingness as 'getting things right' and supports treating fittingness as a distinct normative category.","marker":"[29]"},{"why":"Supplies the explanatory, justificatory, and guiding roles of reasons, along with the 'in a respect' qualifier.","marker":"[51]"},{"why":"Gives the account of normative explanation and justification used to require reason-based explanations of overall moral categories.","marker":"[46]"},{"why":"Presents the RL approach with a single 'true ethical utility function' that the paper analyzes as making moral categories meaningful but moral behavior incidental.","marker":"[1]"},{"why":"Presents the constraint-based policy orchestration approach that the paper argues leaves no principled basis for moral classification.","marker":"[32]"},{"why":"Presents the multi-objective RL approach with a vector of ethical values that the paper judges structurally well-suited but limited by total ordering and the placement of achievement.","marker":"[36]"}],"fun_headline_variants":["Moral RL requires a dedicated reward function","Without a moral reward, RL agents stay amoral","A moral RL agent needs a moral reward function","No moral reward signal, no moral RL agent"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework hinges on the assumption, stated in Section 5 rather than argued, that a moral domain in an RL system must be represented by a dedicated reward function; if morality could instead be encoded as constraints on a non-moral MDP, as learned human preferences without a separate reward channel, or as a safety layer, the paper's criterion for moral classification would fail.","fun_headline_variants_meta":{"raw":{"variants":["Moral RL requires a dedicated reward function","Without a moral reward, RL agents stay amoral","A moral RL agent needs a moral reward function","No moral reward signal, no moral RL agent"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001083,"raw_usage":{"total_tokens":4490,"prompt_tokens":868,"completion_tokens":3622,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":484,"completion_tokens_details":{"reasoning_tokens":3563}},"tokens_in":484,"tokens_out":3622,"duration_ms":24888,"temperature":1.0,"reasoning_tokens":3563,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:14:54.961217+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment in which an RL agent trained solely with human feedback, without any separate moral reward function, outperforms a moral-reward-trained agent on a battery of moral scenarios in new environments would directly contradict the claim that a distinct moral reward function is necessary for moral classification.","supporting_citations":[{"cited_title":"In: Rowland, R.A","cited_arxiv_id":null,"evidence_quote":"Supplies the taxonomy of deontic, evaluative, and fittingness categories and the features that distinguish them."},{"cited_title":"In: Lord, E., Maguire, B","cited_arxiv_id":null,"evidence_quote":"Provides the overall versus contributory distinction and the treatment of normative domains and the weight of reasons."},{"cited_title":"Oxford University Press (2023)","cited_arxiv_id":null,"evidence_quote":"Defines fittingness as 'getting things right' and supports treating fittingness as a distinct normative category."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the explanatory, justificatory, and guiding roles of reasons, along with the 'in a respect' qualifier."},{"cited_title":"Noûs55(1), 3–22 (2021)","cited_arxiv_id":null,"evidence_quote":"Gives the account of normative explanation and justification used to require reason-based explanations of overall moral categories."},{"cited_title":"In: AAAI Workshop: AI, Ethics, and Society (2016) Metanormative Theory for RL-Based Moral Agents 19","cited_arxiv_id":null,"evidence_quote":"Presents the RL approach with a single 'true ethical utility function' that the paper analyzes as making moral categories meaningful but moral behavior incidental."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Presents the constraint-based policy orchestration approach that the paper argues leaves no principled basis for moral classification."},{"cited_title":"Artificial Intelligence (2026)","cited_arxiv_id":null,"evidence_quote":"Presents the multi-objective RL approach with a vector of ethical values that the paper judges structurally well-suited but limited by total ordering and the placement of achievement."}],"review_version":1}