Pith. sign in

REVIEW 4 major objections 6 minor 52 references

Metanormative Theory for RL-Based Moral Agents

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read An RL agent's behavior counts as moral only when a distinct reward function grounds a moral domain.

desk verdict A clear, useful mapping of metanormative categories onto RL components, but the core criterion that a moral domain requires a dedicated reward function is asserted rather than argued, and it drives a negative verdict on constrained-RL approaches that overreaches. read the letter →

arxiv 2608.08220 v1 pith:BCQSDJY5 submitted 2026-08-08 cs.AI

classification cs.AI
keywords machineethicsmetanormativetheoryreinforcementlearningvaluealignmentmoralrewardfunctionnormativecategoriesRLHFmulti-objectiveRL
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper brings tools from metanormative theory—the branch of philosophy that studies the structure of normative domains—to bear on reinforcement-learning (RL) agents. It argues that an RL agent's behavior can be classified as moral only when a distinct reward function establishes a moral domain within the architecture, so that states, actions, and policies can be assessed as morally right, wrong, good, or bad against that standard. The paper derives this criterion by mapping the standard components of an RL system onto four families of normative categories: deontic, evaluative, fittingness, and reason-based. It then applies the criterion to three prominent RL ethics proposals, finding that one makes moral talk meaningful but treats moral behavior as incidental, another leaves no principled basis for moral classification, and a third is structurally well-suited but misplaces the prudential value of achievement and imposes a context-invariant ordering of moral reasons. If the framework is right, it gives researchers a principled way to say when a trained RL policy is genuinely moral rather than merely harm-avoiding.

What carries the argument

The machinery that carries the argument is the pairing of a dedicated reward function with the taxonomy of normative categories from metanormative theory. The reward function—conceived as the 'basis for the moral domain'—supplies the standard against which normative categories apply, while the categories (deontic, evaluative, fittingness, and reason-based) provide the vocabulary used to qualify the components of an RL system: state-action pairs, policies, and reward signals. This mapping, summarized in the paper's Table 2, does the work of separating the moral domain from the prudential domain and yields the paper's evaluative verdicts on concrete RL proposals.

What would settle it

A controlled experiment in which an RL agent trained solely with human feedback, without any separate moral reward function, outperforms a moral-reward-trained agent on a battery of moral scenarios in new environments would directly contradict the claim that a distinct moral reward function is necessary for moral classification.

Watch

Extended reading notes

Core claim

The paper's central claim is that moral vocabulary applied to RL agents is legitimate only when the architecture contains a distinct moral reward function that grounds a moral domain separate from the domain of self-interest or prudence. Given such a function, policies can be morally best, right, and fitting, and state-action pairs can be morally good or bad; without it, there is no principled basis for applying moral categories to the system's behavior. The paper therefore reads the moral status of an RL agent's behavior as determined by the moral status of its states, actions, reward signals, learning algorithms, and policies, and it treats the deontic, evaluative, fittingness, and reason-based categories from metanormative theory as the vocabulary that makes that determination precise. On this view, a policy is moral when it maximizes the moral reward signal, and talk of 'moral behavior' is derivative from the morally optimal policy, just as 'rational behavior' is derivative from the reward-maximizing policy in the prudential domain.

Load-bearing premise

The framework hinges on the assumption, stated in Section 5 rather than argued, that a moral domain in an RL system must be represented by a dedicated reward function; if morality could instead be encoded as constraints on a non-moral MDP, as learned human preferences without a separate reward channel, or as a safety layer, the paper's criterion for moral classification would fail.

Editorial extensions

If this is right

  • If the framework is correct, RLHF-based agents that learn only from human preference labels, with no separate moral reward function, cannot be classified as moral under a principled criterion.
  • Constraint-based approaches that blend domain rewards with ethical penalties produce policies that belong to neither the moral nor the rational domain, so calling them 'ethical' is ungrounded.
  • Multi-objective RL with a vector of ethical reward functions is structurally compatible with moral categories, but only if the value of achievement is placed in the prudential domain and the relative weight of moral reasons is allowed to vary across contexts.
  • An RL policy learned by maximizing expected reward is not thereby moral; the reward function being optimized must be the moral reward function for moral vocabulary to apply.
  • Value talk in AI becomes precise only when translated into normative categories; otherwise claims about 'values' in alignment research lack a clear normative basis.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit a testable design principle: two behaviorally identical policies could differ in moral status if one maximizes a moral reward and the other merely happens to avoid harm while maximizing a domain reward.
  • The paper's critique of scalar rewards as 'mere numbers' suggests that moral RL may need richer reward representations, such as vector-valued rewards with a lexicographic ordering, to capture different kinds of moral reasons.
  • The framework could be pressed on safety-layer and action-masking approaches, where moral constraints are implemented without any reward channel; the paper does not say whether such a filter would count as a moral domain, but its criterion suggests it would not.
  • The criterion could be operationalized as an audit: inspect the reward function of a deployed RL system and check whether a distinct moral reward channel exists before attributing moral status to its behavior.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper proposes a metanormative-theoretic framework for classifying and evaluating RL-based moral agents. It introduces normative categories from metanormative theory (deontic, evaluative, fittingness, and reason-based), distinguishes normative domains, and maps these categories onto components of RL systems (state-action pairs, policies, and reward signals). The central thesis is that moral vocabulary in RL contexts is legitimate only when a distinct moral reward function grounds a moral domain. The authors apply the framework to three existing approaches: Abel et al.'s POMDP-based ethical decision making, Noothigattu et al.'s policy orchestration with learned constraints, and Rodriguez-Soto et al.'s multi-objective RL. They argue that Abel et al. and Rodriguez-Soto et al. provide bases for moral categories (with weaknesses), while Noothigattu et al. do not.

Significance. If the framework were accepted, it would give AI researchers a principled vocabulary for distinguishing moral from merely goal-directed RL systems and a basis for comparing approaches. The paper's strength is its clear synthesis of a substantial philosophical literature and its concrete mapping in Table 2, which is a useful conceptual tool. The authors are also explicit about the simplified and contested nature of the philosophy. However, the central criterion—the requirement of a dedicated moral reward function—is stipulated rather than derived, and the negative verdicts on Abel et al. and Noothigattu et al. depend on this stipulation. The paper is therefore a promising starting point rather than a settled framework.

major comments (4)
  1. [Section 5] Section 5, paragraph beginning 'How should one represent the moral domain in an RL system?': The claim that 'the first thing one needs is a distinct reward function that can serve as a basis for the moral domain' is presented as 'the most natural and conservative answer' but is not derived from the metanormative theory presented in Sections 3–4. In that theory, a normative domain is fixed by a standard against which deontic, evaluative, and fittingness categories apply; nothing in the account requires that standard to be encoded as a scalar reward function. A constrained MDP with hard prohibitions supplies a standard: violated actions are forbidden, compliant policies are permissible, and the constrained optimum is fitting. The paper therefore needs either to argue that only reward functions can serve as such standards, or to restrict the criterion to RL systems that use reward-based normative encodings. As written, the criterion is an engineering choice, and the later negative verdict on Noothigattu et al. in Section 6 inherits this presupposition.
  2. [Section 6, Abel et al.] The claim that the learned policy is moral only 'incidentally' appears to contradict the paper's own Section 5 mapping. If the POMDP's 'true ethical utility function' is the reward function that grounds the moral domain, then the optimal policy is, by the Section 5 definition, the morally best/right/fitting policy, and maximizing expected reward with respect to that function just is maximizing moral reward. The sentence 'the question of whether a policy is morally good, right, or fitting has little to do with maximizing expected reward signal (as opposed to maximizing moral reward signal)' conflates the two, because in this example the expected reward is the expected moral reward. The authors should clarify whether they take the 'true ethical utility function' to be the reward function actually used in optimization, or whether they intend a distinction between observable reward and underlying utility; the current text does not support the 'incidental' criticism.
  3. [Section 6, Noothigattu et al.] The statement that 'there is no principled basis for the application of moral categories in this context' overreaches. The learned constraint rewards in the πC policy define a normative standard: eating ghosts is assigned a strongly negative value, which can ground the deontic judgment that this action is forbidden within the constrained domain, and the constrained optimum can be called fitting. The problem the authors identify is better described as a blending of two domains (domain reward and moral penalty) that obscures which standard is in force, not as the absence of any basis for moral categories. This reformulation would preserve the criticism without making it depend on the unargued reward-function requirement.
  4. [Section 2 and Table 2] The paper states that the moral status of an RL agent's behavior 'must ultimately depend on the moral status of some of the following: states; actions; reward signals; learning algorithms; and policies,' but the subsequent analysis in Section 5 and Table 2 only covers state-action pairs, policies, and reward signals; learning algorithms are never assigned normative categories. Since the list is presented as exhaustive of the possible loci of moral status, the omission leaves the framework incomplete with respect to its own central claim. The authors should either extend the analysis to learning algorithms or remove them from the list.
minor comments (6)
  1. [Section 2] The word 'utalitarianism' should be 'utilitarianism'; similarly, 'POMPD' in Section 6 should be 'POMDP'.
  2. [Section 3] The sentence 'If normatively sensitive situations ar thought of' should read 'are thought of'.
  3. [Section 6] The sentence 'What is unusual.' is incomplete and appears to be a leftover from an intended contrast; the authors should complete or delete it.
  4. [Tables 1 and 3] The table captions read 'T able' instead of 'Table'; please fix the formatting.
  5. [Section 6, Table 3] The entry for Abel et al. lists 'the agent's sole goal is to act morally' as a weakness, but the surrounding text criticizes the agent for learning only one of the morally optimal policies; spelling out why having only this goal is a weakness would make the table more informative.
  6. [Section 7] The concluding caveat that the criterion 'should not be decisive' sits uneasily with the categorical claims in Section 6, such as 'there is no principled basis for the application of moral categories'; the authors should reconcile the hedged conclusion with the body's categorical verdicts.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the framework applies metanormative categories by explicit stipulation, and the paper's self-citations are background only.

full rationale

The paper does not derive empirical predictions or fitted results; it offers a conceptual mapping from metanormative theory to RL components. The central requirement that a moral domain needs a distinct reward function is explicitly presented as 'the most natural and conservative answer' in Section 5, not as a theorem of metanormative theory. Consequently, the negative verdict on Noothigattu et al. is conditional on that stipulation rather than a circular reduction of a derived claim to its own inputs. The discussion of Abel et al. contains an internal tension (if the 'true ethical utility function' is the POMDP reward, then maximizing expected reward is maximizing moral reward), but this is an inconsistency in the critique, not a circularity in the derivation. The self-citations ([2], [30], [48]) are background surveys and examples and do not carry the argument. No equations, fitted parameters, or imported uniqueness theorems are used to force the conclusions, so the paper is best rated as essentially non-circular.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces no fitted constants or new entities. Its central framework rests on five philosophical assumptions: the acceptance of a particular normative taxonomy, the bottoming-out requirement for moral vocabulary, the independence of moral and prudential domains, the representational claim that a dedicated reward function is the right home for a normative domain in RL, and the preference for contributory explanations. The weakest is the reward-function representation claim because the case-study verdicts depend on it.

assumptions (5)
  • domain assumption The metanormative taxonomy of four classes of normative categories (deontic, evaluative, fittingness, reason-based) is the right frame for analyzing moral behavior in AI agents.
    Section 3 presents this as 'a fairly standard take' while noting 'much disagreement among metanormative theorists'; the paper does not defend the taxonomy against alternatives.
  • domain assumption Moral vocabulary such as 'moral' and 'ethical' must 'bottom out' in moral categories pertaining to the moral domain.
    Section 4 concluding bullets state this as a requirement; it is a philosophical stipulation about the semantics of moral terms.
  • domain assumption The moral domain is an independent normative domain, separate from prudence or self-interest.
    Section 4 introduces different normative domains and treats the moral versus prudential contrast as standard; the case evaluation depends on keeping these domains distinct.
  • ad hoc to paper A distinct reward function is the correct representation of a normative domain in an RL system.
    Section 5: 'the most natural and conservative answer' is a dedicated reward function; no argument rules out constraint-based or preference-based encodings.
  • domain assumption Approaches that can explain overall normative categories in terms of contributory (reason-based) categories are to be preferred.
    Section 4 and Conclusion endorse this 'division of labor' in normative explanation; the paper explicitly says this criterion is not decisive.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Metanormative Theory for RL-Based Moral Agents." pith.science (2026). https://pith.science/paper/BCQSDJY5

@misc{pith2026260808220,
  author       = {Pith},
  title        = {Pith review of: Metanormative Theory for RL-Based Moral Agents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BCQSDJY5}},
  note         = {Machine review of arXiv:2608.08220}
}
read the original abstract

The overlapping disciplines of machine ethics and value alignment are concerned with designing artificial agents that are aligned with human values and that act in ethically acceptable ways. A recent trend in these disciplines is the use of reinforcement learning (RL) to design such agents, sidelining the philosophical literature that used to play a more central role. Against this backdrop, this paper pursues two goals. The first is to draw out ideas from recent work in metanormative theory that can be useful for designing artificial moral and value-aligned agents. The second is to examine the RL architecture through the lens of these ideas. This will give us clearer criteria for when an RL agent's behavior can be classified as moral, as well as a basis for evaluating and comparing different RL-based approaches to machine ethics and value alignment.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 47 canonical work pages

  1. [1]

    In: AAAI Workshop: AI, Ethics, and Society (2016) Metanormative Theory for RL-Based Moral Agents 19

    Abel, D., MacGlashan, J., Littman, M.L.: Reinforcement learning as a framework for ethical decision making. In: AAAI Workshop: AI, Ethics, and Society (2016) Metanormative Theory for RL-Based Moral Agents 19

  2. [2]

    In: Proceedings of the Seventh AAAI/ACM Conference on AI, Ethics, and Society (AIES-2024)

    Alcaraz, B., Knoks, A., Streit, D.: Estimating weights of reasons using meta- heuristics: A hybrid approach to machine ethics. In: Proceedings of the Seventh AAAI/ACM Conference on AI, Ethics, and Society (AIES-2024). pp. 27–38. ACM Press (2024)

  3. [3]

    Minds and Machines17(1), 1–10 (2007)

    Anderson, M., Anderson, S.L.: The status of machine ethics: A report from the AAAI Symposium. Minds and Machines17(1), 1–10 (2007)

  4. [4]

    (eds.): Machine Ethics

    Anderson, M., Anderson, S.L. (eds.): Machine Ethics. Cambridge University Press (2011)

  5. [5]

    Routledge (2017)

    Baker,D.:Thevarietiesofnormativity.In:TheRoutledgeHandbookofMetaethics. Routledge (2017)

  6. [6]

    In: Shafer-Landau, R

    Baker, D.: Skepticism about oughtsimpliciter. In: Shafer-Landau, R. (ed.) Oxford Studies in Metaethics 13. Oxford University Press (2018)

  7. [7]

    In: Sun, R

    Bello, P., Malle, B.F.: Computational approaches to morality. In: Sun, R. (ed.) Cambridge Handbook of Computational Cognitive Sciences, pp. 1037–1063. Cam- bridge University Press (2023)

  8. [8]

    Ethics118(1), 109–139 (2007)

    Berker, S.: Particular reasons. Ethics118(1), 109–139 (2007)

Show all 52 references
  1. [9]

    In: Rowland, R.A

    Berker, S.: The deontic, the evaluative, and the fitting. In: Rowland, R.A. (ed.) Fit- tingness: Essays in the Philosophy of Normativity. Oxford University Press (2022)

  2. [10]

    In: IJCAI-17 Workshop on Explainable AI (XAI)

    Biran, O., Cotton, C.: Explanation and justification in machine learning: A survey. In: IJCAI-17 Workshop on Explainable AI (XAI). vol. 8, pp. 8–13 (2017)

  3. [11]

    Australasian Journal of Philos- ophy102(2), 497–511 (2024)

    Brown, J.: On scepticism about ought simpliciter. Australasian Journal of Philos- ophy102(2), 497–511 (2024)

  4. [12]

    In: Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (AIES-2022)

    Canavotto,I.,Horty,J.:Piecemealknowledgeacquisitionforcomputationalnorma- tive reasoning. In: Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (AIES-2022). pp. 171–80. ACM Press (2022)

  5. [13]

    In: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency

    Chan, A., Salganik, R., Markelius, A., Pang, C., Rajkumar, N., Krasheninnikov, D., et al.: Harms from increasingly agentic algorithmic systems. In: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. p. 651–666. FAccT ’23, Association for Comp...

  6. [14]

    Conitzer, V.: Why should we ever automate moral decision making? (2024), https://arxiv.org/abs/2407.07671

  7. [15]

    arXiv preprint arXiv:2404.10271 (2024)

    Conitzer, V., Freedman, R., Heitzig, J., Holliday, W.H., Jacobs, B.M., Lambert, N., et al.: Social choice for AI alignment: Dealing with diverse human feedback. arXiv preprint arXiv:2404.10271 (2024)

  8. [16]

    Blackwell (1993)

    Dancy, J.: Moral Reasons. Blackwell (1993)

  9. [17]

    In: Belousov, B., Abdulsamad, H., Klink, P., Parisi, S., Peters, J

    Eschmann, J.: Reward function design in reinforcement learning. In: Belousov, B., Abdulsamad, H., Klink, P., Parisi, S., Peters, J. (eds.) Reinforcement Learning Algorithms: Analysis and Applications, pp. 25–33. Springer (2021)

  10. [18]

    AI Ethics6(77) (2026)

    Faroldi, F.L.G.: Reasons-based artificial agents. AI Ethics6(77) (2026)

  11. [19]

    Minds and Machines 14(3), 349–379 (2004)

    Floridi, L., Sanders, J.: On the morality of artificial agents. Minds and Machines 14(3), 349–379 (2004)

  12. [20]

    Minds and Machines 30(3), 411–437 (2020)

    Gabriel, I.: Artificial Intelligence, values, and alignment. Minds and Machines 30(3), 411–437 (2020)

  13. [21]

    The MIT Press (2021)

    Hidalgo, C.A., Orghian, D., Canals, J.A., de Almeida, F., Martin, N.: How Humans Judge Machines. The MIT Press (2021)

  14. [22]

    In: Howard, C., Cosker-Rowland, R

    Howard, C., Cosker-Rowland, R.: Fittingness: A user’s guide. In: Howard, C., Cosker-Rowland, R. (eds.) Fittingness. Oxford University Press (2022)

  15. [23]

    Philosophy Compass13(11), e12542 (2018)

    Howard, C.: Fittingness. Philosophy Compass13(11), e12542 (2018)

  16. [24]

    Knoks and M

    Ji, J., Qiu, T., Chen, B., Zhang, B., Lou, H., Wang, K., et al.: AI alignment: A comprehensive survey (2025), http://arxiv.org/abs/2310.19852, arXiv:2310.19852 20 A. Knoks and M. Slavkovik

  17. [25]

    Westview Press (1998)

    Kagan, S.: Normative Ethics. Westview Press (1998)

  18. [26]

    Mind134(536), 1164–1173 (2025)

    Kauppinen, A.: Review ofGetting Things Right: Fittingness, Reasons, and Value, by Conor McHugh and Jonathan WayThe Range of Reasons in Ethics and Epis- temology, by Daniel Whiting. Mind134(536), 1164–1173 (2025)

  19. [27]

    In: Lord, E., Maguire, B

    Lord, E., Maguire, B.: An opinionated guide to the weight of reasons. In: Lord, E., Maguire, B. (eds.) Weighing Reasons, pp. 3–24. Oxford University Press (2016)

  20. [28]

    Ethics126(3), 575–606 (2016)

    McHugh, C., Way, J.: Fittingness first. Ethics126(3), 575–606 (2016)

  21. [29]

    Oxford University Press (2023)

    McHugh, C., Way, J.: Getting Things Right: Fittingness, Reasons, and Value. Oxford University Press (2023)

  22. [30]

    Montes, N., Osman, N., Sierra, C., Slavkovik, M.: Value engineering for au- tonomous agents (2023), https://arxiv.org/abs/2302.08759

  23. [31]

    In: Proceed- ings of the Seventeenth International Conference on Machine Learning

    Ng, A.Y., Russell, S.J.: Algorithms for inverse reinforcement learning. In: Proceed- ings of the Seventeenth International Conference on Machine Learning. p. 663–670. ICML ’00, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA (2000)

  24. [32]

    Noothigattu, R., Bouneffouf, D., Mattei, N., Chandra, R., Madan, P., Varshney, K.R., et al.: Teaching AI agents ethical values using reinforcement learning and policyorchestration.In:ProceedingsoftheTwenty-EighthInternationalJointCon- ference on Artificial Intelligence, IJCAI-...

  25. [33]

    Oxford University Press (2011)

    Parfit, D.: On What Matters (Volume I). Oxford University Press (2011)

  26. [34]

    Oxford University Press (1990)

    Raz, J.: Practical reason and norms. Oxford University Press (1990)

  27. [35]

    Neural Computing and Applications37, 25619–25644 (2023)

    Rodriguez-Soto, M., Lopez-Sanchez, M., Rodriguez-Aguilar, J.A.: Multi-objective reinforcement learning for designing ethical multi-agent environments. Neural Computing and Applications37, 25619–25644 (2023)

  28. [36]

    Artificial Intelligence (2026)

    Rodriguez-Soto, M., Rădulescu, R., Bistaffa, F., Ricart, O., Mayoral, A., Lopez- Sanchez, M., et al.: Multi-objective reinforcement learning for provably incentivis- ing alignment with value systems. Artificial Intelligence (2026)

  29. [37]

    Ethics and Information Technology24(9) (2022)

    Rodriguez-Soto, M., Serramia, M., Lopez-Sanchez, M., Rodriguez-Aguilar, J.A.: Instilling moral value alignment by means of multi-objective reinforcement learn- ing. Ethics and Information Technology24(9) (2022)

  30. [38]

    Pearson Edu- cation, 4 edn

    Russell, S., Norvig, P.: Artificial Intelligence: A Modern Approach. Pearson Edu- cation, 4 edn. (2020)

  31. [39]

    In: Zalta, E.N., Nodelman, U

    Sayre-McCord, G.: Metaethics. In: Zalta, E.N., Nodelman, U. (eds.) The Stan- ford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University, Spring 2023 edn. (2023)

  32. [40]

    Cambridge, MA: Harvard University Press (1998)

    Scanlon, T.M.: What We Owe to Each Other. Cambridge, MA: Harvard University Press (1998)

  33. [41]

    Oxford University Press (2021)

    Schroeder, M.: Reasons First. Oxford University Press (2021)

  34. [42]

    Oxford University Press (2012)

    Shafer-Landau, R.: The Fundamentals of Ethics. Oxford University Press (2012)

  35. [43]

    Philosophy and Public Affairs1(3), 229–43 (1972)

    Singer, P.: Famine, affluence, and morality. Philosophy and Public Affairs1(3), 229–43 (1972)

  36. [44]

    A Bradford Book, Cambridge, MA, USA (2018)

    Sutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction. A Bradford Book, Cambridge, MA, USA (2018)

  37. [45]

    In: Zalta, E.N., Nodelman, U

    Tucker, C.: Weighing reasons. In: Zalta, E.N., Nodelman, U. (eds.) The Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University, Win- ter 2025 edn. (2025)

  38. [46]

    Noûs55(1), 3–22 (2021)

    Väyrynen, P.: Normative explanation and justification. Noûs55(1), 3–22 (2021)

  39. [47]

    In: Copp, D., Rosati, C

    Väyrynen, P.: Varieties of normative explanation. In: Copp, D., Rosati, C. (eds.) The Oxford Handbook of Metaethics. Oxford University Press (forthcoming)

  40. [48]

    Vishwanath, A., Dennis, L.A., Slavkovik, M.: Reinforcement learning and machine ethics: A systematic review (2024), https://arxiv.org/abs/2407.02425 Metanormative Theory for RL-Based Moral Agents 21

  41. [49]

    Oxford University Press, Inc., USA (2008)

    Wallach, W., Allen, C.: Moral Machines: Teaching Robots Right from Wrong. Oxford University Press, Inc., USA (2008)

  42. [50]

    AI & SOCIETY22(4), 565–582 (2008)

    Wallach, W., Allen, C., Smit, I.: Machine morality: Bottom-up and top-down ap- proaches for modelling human moral faculties. AI & SOCIETY22(4), 565–582 (2008)

  43. [51]

    Whiting,D.:TheRangeofReasons:InEthicsandEpistemology.OxfordUniversity Press (2022)

  44. [52]

    Journal of Artificial Intelligence Research82(2025)

    Zhong, T., Song, Y., Limarga, R., Pagnucco, M.: Computational machine ethics: A survey. Journal of Artificial Intelligence Research82(2025)

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.