Pith. sign in

REVIEW 3 major objections 6 minor 50 references

Moral theories are resource strategies on a breadth–depth tradeoff, not rival truths.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 14:33 UTC pith:NIKJUCXA

load-bearing objection Clean formalization of moral breadth vs. depth with a usable regret notion; the neutrality claim is overstated because M* is consequentialist scaffolding. the 3 major comments →

arxiv 2607.00002 v1 pith:NIKJUCXA submitted 2026-04-01 cs.AI cs.CY

Bounded Morality: Defining the Space of Moral Computation

classification cs.AI cs.CY
keywords bounded rationalityresource rationalitymoral agentsethical theoriesAI alignmentmoral progressmoral breadthmoral depth
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper claims that moral reasoning for any finite agent is a constrained computational problem whose demand structure can be mapped along two axes: moral breadth (how many entities, groups, and timescales are treated as relevant) and moral depth (how far interactions and consequences are integrated by inference). Limited attention, memory, time, and compute force an unavoidable tradeoff between those axes, so only a subset of the breadth–depth plane is feasible. Within that feasible region, familiar ethical theories—utilitarianism, deontology, contractualism, virtue ethics, care ethics—are reinterpreted as locally efficient allocation strategies suited to different demand regimes, not as competing accounts of moral truth. The paper defines moral regret relative to an unbounded ideal and moral progress as either better allocation under a fixed budget or expansion of the frontier itself. For artificial systems, the implication is that alignment should scale and allocate moral reasoning capacity rather than copy human judgments that already encode human bounds.

Core claim

Bounded Morality holds that finite agents face a structural tradeoff between moral breadth and moral depth; ethical theories are characteristic resource-allocation strategies on that frontier, and moral regret and progress are well-defined relative to an unbounded ideal under fixed computational budgets.

What carries the argument

The breadth–depth frontier under a resource budget: abstract moral graphs of breadth b(G)=|V|+|E| and rollout depth H, with cost Cost=αb+βbH^p, inducing a downward-sloping Pareto frontier and strategy-relative moral regret against unbounded M*.

Load-bearing premise

The paper defines ground-truth moral value as infinite-horizon discounted welfare that adds up over people and their pairwise links, while still claiming not to favor any particular moral theory.

What would settle it

Find a realistic moral domain where agents who expand both breadth and depth under larger budgets do not reduce measured regret relative to a fixed welfare objective, or where the main ethical theories do not occupy distinct, locally efficient regions of the breadth–depth plane.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes Bounded Morality, a computational-level framework that treats moral reasoning by finite agents as constrained inference over a graph-structured moral system. Moral situations are characterized along two axes—moral breadth (representational scope of entities and interactions) and moral depth (inferential horizon)—whose joint cost under a fixed budget yields a Pareto frontier of feasible allocations. Ethical theories are reinterpreted as families of resource-bounded strategies that occupy different regions of this space; moral regret and moral progress are defined relative to an unbounded ideal; and AI alignment is argued to depend on scaling and allocating moral reasoning capacity rather than imitating human judgments. A content-moderation example illustrates depth reversal and breadth-induced infeasibility.

Significance. If the framing holds, the paper offers a useful unifying vocabulary for moral disagreement, heuristics, and theory choice under constraint, and a concrete alternative to pure imitation-based alignment. Strengths include an internally consistent formalization (Defs. 3.1–3.17), elementary but correctly stated cost lower bounds and inverse scaling (Props. 3.19–3.20, Cor. 3.22), and a fully worked numerical example (Appendix B) that makes depth/breadth regret operational. The contribution is primarily conceptual and diagnostic rather than predictive or empirically fitted; its value for machine ethics and AI alignment depends on whether the theory-neutrality and strategy-efficiency claims can be made precise without smuggling a single ideal moral objective.

major comments (3)
  1. [§3 Defs. 3.3–3.4, 3.15–3.17; §4; Abstract/§1] Defs. 3.3–3.4 and 3.15–3.17 make ground-truth moral value M* an infinite-horizon discounted additive welfare over node/edge potentials. All notions of regret, distributional efficiency, and moral progress are defined as distance to this M*. Section 1 and the abstract simultaneously claim theory-neutrality and that ethical theories are strategies rather than competing accounts of moral truth. Non-consequentialist theories in §4 are then scored by the same additive M* after being recast as restrictions on (G,H), T, or admissibility (Table 2). This is load-bearing: if the ideal is itself a consequentialist commitment, the “strategies not truth” reinterpretation and the alignment implication are relative to that commitment. Either generalize M* to an arbitrary fixed moral evaluation functional (and show regret/progress still work), or explicitly drop theory-neutrality and state that the fram
  2. [§4, Table 2, Figure 2] The central claim that ethical theories are “locally efficient strategies adapted to different demand regimes” (Abstract, §4.7, Fig. 2) is not demonstrated. Section 4 assigns each theory a characteristic (b,H) region and aggregation/admissibility form, but never shows that these allocations minimize expected regret under any (P,B), nor that they dominate alternatives in their regimes. Table 2 and Fig. 2 are labeled conceptual caricatures. Either provide a formal efficiency result (even for a stylized P and cost model) or soften the claim to “structurally distinct allocation families” without efficiency language.
  3. [§1 Implications; §5; Conclusion] The AI-alignment implication—that alignment should scale and allocate moral capacity rather than imitate human judgments—rests on treating human judgments as suboptimal approximations to M*. The paper does not address how M* (or its primitives U, F, G*) is known or elicited for artificial systems when human judgments are the primary evidence. Without a path from observed judgments to the ideal, capacity scaling alone does not specify what is being optimized. A short discussion of identification or multi-objective/ideal-agnostic evaluation would make the implication defensible.
minor comments (6)
  1. [§3.6] Props. 3.19–3.20 are immediate encoding/simulation lower bounds; the write-up can note this more plainly so readers do not expect non-trivial complexity results.
  2. [Def. 3.21, Cor. 3.22] Canonical cost Cost(b,H)=αb+βbH^p introduces free exponents α,β,p with no sensitivity analysis; a brief note on how the qualitative inverse scaling depends on p would help.
  3. [Appendix B] Appendix B parameters (β=0.9, α=0.2, δ=0.05, etc.) are hand-chosen; state explicitly that the example is illustrative, not a calibrated case study.
  4. [Figure 1] Figure 1’s left panel is hard to read (overlapping labels for expected vs state-dependent regret); consider separating panels or simplifying the annotation.
  5. [§1] Related work on resource-rational contractualism (Levine et al., cited) is close in spirit; a sharper contrast paragraph would clarify novelty relative to that line.
  6. [§3] Notation: 𝒜_H vs 𝒜^H and π_V vs π_V appear inconsistently typeset in places; unify.

Circularity Check

1 steps flagged

Framework is stipulative and self-contained; no fitted predictions or load-bearing self-citation chains. Only mild by-construction flavor in the theories-as-strategies mapping.

specific steps
  1. self definitional [§4 (esp. 4.1–4.7, Table 2); Abstract claim that theories “correspond to” strategies]
    "Within this space, ethical theories correspond to locally efficient strategies adapted to different demand regimes rather than competing accounts of moral truth. … In this section, we reinterpret several canonical ethical theories as families of bounded moral strategies. Each theory is modeled as a structured restriction on: 1. admissible allocation rules ρ … The mappings below are deliberately schematic."

    The paper defines strategy families (ℛUTIL prioritizes breadth, ℛCONTRACT prioritizes depth, ℛVIRTUE sets H=0, etc.) so that each matches a named ethical theory, then asserts that ethical theories “correspond to” those strategies. The correspondence is true by construction of the mapping rather than independently derived. Mild only: the paper presents this as reinterpretation/schematic modeling, not as an empirical or uniqueness result.

full rationale

Bounded Morality is a definitional/formal framework paper, not an empirical derivation. Breadth b(G)=|V|+|E|, depth H, Cost(b,H)=αb+βbH^p, regret R=M*−M*(πρ), and moral progress as lower expected regret are introduced by definition (Defs. 3.5–3.17, 3.21) and then used consistently; the inverse scaling Hmax(b)≤C b^{−1/p} (Cor. 3.22) is a direct algebraic consequence of that cost model, not a prediction forced by fitting. The content-moderation example (Appendix B) uses hand-chosen parameters and is explicitly illustrative. There are no self-citations of uniqueness theorems or prior results by the same authors that force the central claims; external citations (Simon, Marr, Crimston, Kohlberg, Scanlon, etc.) supply background, not closed loops. The only mild self-definitional flavor is the §4 reinterpretation of ethical theories as strategy families ℛUTIL, ℛRULE, …: the paper constructs allocation rules and aggregation functionals to match each theory, then states that theories “correspond to” those strategies. That correspondence is by modeling choice (and the paper labels the mappings “deliberately schematic”), not a claimed first-principles derivation of an independent fact. That does not rise to circular prediction or load-bearing self-citation. Score 1 for that minor by-construction interpretive step; central formal content is independent and non-circular.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 3 invented entities

The central claims rest on a suite of modeling choices that turn moral evaluation into constrained optimal control on graphs. Most are domain assumptions imported from control, social choice, and welfare economics; a few (the exact cost functional, the claim of orthogonality, the reduction of ethical theories to allocation families) are paper-specific. No free parameters are fitted to external data for the main claims; the example parameters are illustrative only.

free parameters (2)
  • canonical cost exponents/coefficients α, β, p
    Chosen for tractability in Def. 3.21 to produce the inverse scaling H_max(b)=Θ(b^{-1/p}); not derived from first principles or measured.
  • example dynamics parameters (β=0.9, α=0.2, δ=0.05, r0=1.0, k=0.1, γ=0.9, λ=0.1)
    Hand-set in Appendix B to produce a clean depth-reversal demonstration; not fitted to real moderation data.
axioms (5)
  • domain assumption Moral value is infinite-horizon discounted graph-structured welfare decomposable into node and edge potentials (Defs. 3.3–3.4).
    Imported from welfare economics / graphical models; load-bearing for the definition of regret and the claim of theory-neutrality.
  • domain assumption Dynamics are local on the moral interaction graph (Def. 3.2).
    Standard locality assumption from graphical models and multi-agent systems; enables the Θ(Hb) simulation cost.
  • standard math Informational cost is strictly increasing in breadth b(G)=|V|+|E| and inferential cost scales at least as Hb (Props. 3.19–3.20).
    Elementary counting lower bounds once the representation and rollout are defined.
  • ad hoc to paper Breadth and depth are the two primary orthogonal demand dimensions of moral situations.
    Motivated by developmental and dual-process literature but stipulated as the core axes of the framework (§2).
  • ad hoc to paper Ethical theories can be adequately caricatured as structured restrictions on allocation rules ρ, aggregation functionals T, and admissibility sets (§4).
    Explicitly labeled schematic; required for the reinterpretation claim.
invented entities (3)
  • Moral breadth b(G) and moral depth H as the two axes of moral computation independent evidence
    purpose: Define the demand structure and the feasible region under resource budgets.
    New formalization; independent psychological evidence for variation along each axis is cited, but the joint computational geometry is introduced here.
  • Bounded moral strategy ρ and moral regret R(ω;ρ,B) no independent evidence
    purpose: Provide a performance measure and a definition of moral progress under fixed budgets.
    Direct analogues of regret in online learning / approximate dynamic programming, specialized to moral graphs.
  • Canonical cost model Cost(b,H)=αb+βbH^p and the resulting inverse scaling law no independent evidence
    purpose: Make the breadth-depth Pareto frontier analytically tractable.
    Chosen for convenience; not independently measured.

pith-pipeline@v1.1.0-grok45 · 22821 in / 3331 out tokens · 34653 ms · 2026-07-13T14:33:06.642598+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Bounded Morality: Defining the Space of Moral Computation." pith.science (2026). https://pith.science/paper/NIKJUCXA

@misc{pith2026260700002,
  author       = {Pith},
  title        = {Pith review of: Bounded Morality: Defining the Space of Moral Computation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NIKJUCXA}},
  note         = {Machine review of arXiv:2607.00002}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Moral cognition has traditionally been modeled as adherence to fixed ethical theories--deontology, consequentialism, virtue ethics--implemented as static rules or value functions. We propose Bounded Morality, a formal framework for analyzing the computational demands of moral problems faced by finite agents. Extending Herbert Simon's notion of bounded rationality, we formalize moral situations along two orthogonal dimensions: moral breadth, the scope of entities treated as morally relevant, and moral depth, the inferential integration required to evaluate their interactions. Limited resources impose an unavoidable tradeoff between these dimensions, defining a feasible space of moral computation. Within this space, ethical theories correspond to locally efficient strategies adapted to different demand regimes rather than competing accounts of moral truth. The framework yields a formal notion of moral regret and moral progress under constraint, and implies that moral alignment in artificial systems depends on the scaling and allocation of moral reasoning capacity rather than on direct imitation of human judgments.

Figures

Figures reproduced from arXiv: 2607.00002 by Caryn Tran, Max Kanwal, Patrick Mineault.

Figure 1
Figure 1. Figure 1: At fixed budget 𝐵 and distribution P, policies 𝜌old and 𝜌new map world states to structural allocations (𝐺, 𝐻 ). The plot highlights a representative state 𝜔0 where the improved policy selects a different breadth–depth trade-off (𝐺, 𝐻 ) = 𝜌new(𝜔0 , 𝐵) on the fixed-budget frontier, yielding lower state-dependent regret 𝑅(𝜔0 ; 𝜌new, 𝐵) than 𝑅(𝜔0 ; 𝜌old, 𝐵). Dashed lines indicate expected regret 𝔼𝜔 [𝑅(𝜔; 𝜌, 𝐵… view at source ↗
Figure 2
Figure 2. Figure 2: Budget contours (𝐵) induce Pareto-efficient breadth–depth frontiers under a canonical cost model. Ethical theories are depicted as strategy families with characteristic scaling tendencies: utilitarianism trades budget for breadth, contractualism for depth, care ethics for local breadth and moderate depth, while deontology and virtue ethics remain near low-depth regimes. Regions are conceptual caricatures. … view at source ↗
Figure 3
Figure 3. Figure 3: Ground-truth moral interaction graph 𝐺 ⋆ (left) and a coarse abstraction 𝐺 (right) induced by aggre￾gation map 𝜋𝑉 with 𝑋 = {𝐸1 , 𝐸2 , 𝐶} and 𝑌 = {𝑀1 , 𝑀2 }. Mechanism of Depth Reversal. With a short horizon, the planner mainly sees the immediate ben￾efit of cutting the bridge between extremists and moderates: sanctioning 𝐶 quickly reduces visible spillover, which appears strongly beneficial in the first fe… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

50 extracted references · 14 canonical work pages

  1. [1]

    Takeshita, R

    M. Takeshita, R. Rafal, K. Araki, Towards Theory-based Moral AI: Moral AI with Aggregating Models Based on Normative Ethical Theory, 2023. URL: http://arxiv.org/abs/2306.11432. doi: 10. 48550/arXiv.2306.11432 , arXiv:2306.11432 [cs]

  2. [2]

    Hegde, V

    A. Hegde, V. Agarwal, S. Rao, Ethics, Prosperity, and Society: Moral Evaluation Using Virtue Ethics and Utilitarianism, in: Proceedings of the Twenty-Ninth International Joint Confer- ence on Artificial Intelligence, International Joint Conferences on Artificial Intelligence Orga- nization, Yokohama, Japan, 2020, pp. 167–174. URL: https://www.ijcai.org/pr...

  3. [3]

    G. R. T. White, A. Samuel, P. Jones, N. Madhavan, A. Afolayan, A. Abdullah, T. Kaushik, Mapping the ethic-theoretical foundations of artificial intelligence re- search, Thunderbird International Business Review 66 (2024) 171–183. URL: https: //onlinelibrary.wiley.com/doi/abs/10.1002/tie.22368. doi: 10.1002/tie.22368 , _eprint: https://onlinelibrary.wiley....

  4. [4]

    Preniqi, I

    V. Preniqi, I. Ghinassi, J. Ive, C. Saitis, K. Kalimeri, MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions, in: Proceedings of the 2024 International Con- ference on Information Technology for Social Good, ACM, Bremen Germany, 2024, pp. 433–442. URL: https://dl.acm.org/doi/10.1145/3677525.3678694. doi:10.1145/3677525.3678694

  5. [5]

    Gabriel, Artificial intelligence, values, and alignment, Minds and machines 30 (2020) 411–437

    I. Gabriel, Artificial intelligence, values, and alignment, Minds and machines 30 (2020) 411–437

  6. [6]

    H. A. Simon, Bounded Rationality, in: J. Eatwell, M. Milgate, P. Newman (Eds.), Utility and Probability, Palgrave Macmillan UK, London, 1990, pp. 15–18. URL: https://doi.org/10.1007/ 978-1-349-20568-4_5 . doi:10.1007/978- 1- 349- 20568- 4_5

  7. [7]

    J. R. Anderson, The Adaptive Character of Thought, 1 ed., Psychology Press, 2013. URL: https: //www.taylorfrancis.com/books/9780203771730. doi:10.4324/9780203771730

  8. [8]

    T. L. Griffiths, F. Lieder, N. D. Goodman, Rational Use of Cognitive Resources: Levels of Analysis Between the Computational and the Algorithmic, Topics in Cognitive Science 7 (2015) 217–229. URL: https://onlinelibrary.wiley.com/doi/10.1111/tops.12142. doi:10.1111/tops.12142

  9. [9]

    Levine, N

    S. Levine, N. Chater, J. B. Tenenbaum, F. Cushman, Resource-rational contractu- alism: A triple theory of moral cognition, Behavioral and Brain Sciences (2024) 1–38. URL: https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/ article/abs/resourcerational-contractualism-a-triple-theory-of-moral-cognition/ 5A567D41A472DBC0D965460966580C74. d...

  10. [10]

    Levine, M

    S. Levine, M. Franklin, T. Zhi-Xuan, S. Y. Guyot, L. Wong, D. Kilov, Y. Choi, J. B. Tenenbaum, N. Goodman, S. Lazar, Resource Rational Contractualism Should Guide AI Alignment, arXiv preprint arXiv:2506.17434 (2025)

  11. [11]

    Chugh, M

    D. Chugh, M. H. Bazerman, M. R. Banaji, Bounded Ethicality as a Psychological Barrier to Rec- ognizing Conflicts of Interest, in: Conflicts of interest: Challenges and solutions in business, law, medicine, and public policy, Cambridge University Press, New York, NY, US, 2005, pp. 74–95. doi:10.1017/CBO9780511610332.006

  12. [12]

    A. E. Tenbrunsel, K. Smith‐Crowe, 13 Ethical Decision Making: Where We’ve Been and Where We’re Going, Academy of Management Annals 2 (2008) 545–607. URL: http://journals.aom.org/ doi/10.5465/19416520802211677. doi:10.5465/19416520802211677

  13. [13]

    Marr, Vision: a computational investigation into the human representation and processing of visual information, MIT Press, Cambridge, Mass, 2010

    D. Marr, Vision: a computational investigation into the human representation and processing of visual information, MIT Press, Cambridge, Mass, 2010

  14. [14]

    Piaget, M

    J. Piaget, M. Cook, The origins of intelligence in children, volume 8, International universities press New York, 1952. Issue: 5

  15. [15]

    Kohlberg, Moral stages and moralization: The cognitive-development approach, Moral de- velopment and behavior: Theory research and social issues (1976) 31–53

    L. Kohlberg, Moral stages and moralization: The cognitive-development approach, Moral de- velopment and behavior: Theory research and social issues (1976) 31–53. Publisher: Rinehart & Winston

  16. [16]

    Eisenberg, R

    N. Eisenberg, R. A. Fabes, T. L. Spinrad, Prosocial Development, in: Handbook of child psychology: Social, emotional, and personality development, Vol. 3, 6th ed, John Wiley & Sons, Inc., Hoboken, NJ, US, 2006, pp. 646–718

  17. [17]

    C. R. Crimston, P. G. Bain, M. J. Hornsey, B. Bastian, Moral expansiveness: Examining variability in the extension of the moral world., Journal of personality and social psychology 111 (2016) 636. Publisher: American Psychological Association

  18. [18]

    Singer, The expanding circle, Clarendon Press Oxford, 1981

    P. Singer, The expanding circle, Clarendon Press Oxford, 1981

  19. [19]

    Parfit, Reasons and Persons, 1 ed., Oxford University PressOxford, 1986

    D. Parfit, Reasons and Persons, 1 ed., Oxford University PressOxford, 1986. URL: https://academic. oup.com/book/12484. doi:10.1093/019824908X.001.0001

  20. [20]

    S. M. Gardiner, A perfect moral storm: The ethical tragedy of climate change, Oxford University Press, 2011

  21. [21]

    Waytz, J

    A. Waytz, J. Cacioppo, N. Epley, Who Sees Human?: The Stability and Importance of Individual Differences in Anthropomorphism, Perspectives on Psychological Science 5 (2010) 219–232. URL: https://journals.sagepub.com/doi/10.1177/1745691610369336. doi: 10. 1177/1745691610369336

  22. [22]

    Haidt, The emotional dog and its rational tail: A social intuitionist approach to moral judg- ment, Psychological Review 108 (2001) 814–834

    J. Haidt, The emotional dog and its rational tail: A social intuitionist approach to moral judg- ment, Psychological Review 108 (2001) 814–834. doi: 10.1037/0033- 295X.108.4.814 , place: US Publisher: American Psychological Association

  23. [23]

    J. D. Greene, L. E. Nystrom, A. D. Engell, J. M. Darley, J. D. Cohen, The Neural Bases of Cognitive Conflict and Control in Moral Judgment, Neuron 44 (2004) 389–400. URL: https://linkinghub. elsevier.com/retrieve/pii/S0896627304006348. doi:10.1016/j.neuron.2004.09.027

  24. [24]

    J. D. Greene, S. A. Morelli, K. Lowenberg, L. E. Nystrom, J. D. Cohen, Cognitive load selectively interferes with utilitarian moral judgment, Cognition 107 (2008) 1144–1154. doi: 10.1016/j. cognition.2007.11.004

  25. [25]

    Conway, B

    P. Conway, B. Gawronski, Deontological and utilitarian inclinations in moral decision making: a process dissociation approach, Journal of Personality and Social Psychology 104 (2013) 216–235. doi:10.1037/a0031021

  26. [26]

    R. L. Selman, The Growth of Interpersonal Understanding: Developmental and Clinical Analyses, Academic Press, 1980

  27. [27]

    Piaget, The moral judgment of the child, Routledge, 2013

    J. Piaget, The moral judgment of the child, Routledge, 2013

  28. [28]

    M. Buon, P. Jacob, E. Loissel, E. Dupoux, A non-mentalistic cause-based heuristic in human social evaluations, Cognition 126 (2013) 149–155. doi: 10.1016/j.cognition.2012.09.006

  29. [29]

    J. W. Martin, M. Buon, F. Cushman, The Effect of Cognitive Load on Intent- Based Moral Judgment, Cognitive Science 45 (2021) e12965. URL: https:// onlinelibrary.wiley.com/doi/abs/10.1111/cogs.12965. doi: 10.1111/cogs.12965 , _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/cogs.12965

  30. [30]

    J. D. Greene, Beyond Point-and-Shoot Morality: Why Cognitive (Neuro)Science Matters for Ethics, Ethics 124 (2014) 695–726. URL: https://www.journals.uchicago.edu/doi/10.1086/675875. doi:10.1086/675875

  31. [31]

    Kahane, J

    G. Kahane, J. A. C. Everett, B. D. Earp, L. Caviola, N. S. Faber, M. J. Crockett, J. Savulescu, Beyond sacrificial harm: A two-dimensional model of utilitarian psychology, Psychological Review 125 (2018) 131–164. doi: 10.1037/rev0000093

  32. [32]

    FETHERSTONHAUGH, P

    D. FETHERSTONHAUGH, P. SLOVIC, S. JOHNSON, J. FRIEDRICH, Insensitivity to the Value of Human Life: A Study of Psychophysical Numbing, Journal of Risk and Uncertainty 14 (1997) 283–300. URL: https://doi.org/10.1023/A:1007744326393. doi:10.1023/A:1007744326393

  33. [33]

    R. S. Suter, R. Hertwig, Time and moral judgment, Cognition 119 (2011) 454–458. doi: 10.1016/ j.cognition.2011.01.018

  34. [34]

    J. M. Paxton, L. Ungar, J. D. Greene, Reflection and reasoning in moral judgment, Cognitive Science 36 (2012) 163–177. doi: 10.1111/j.1551- 6709.2011.01210.x

  35. [35]

    Baron, M

    J. Baron, M. Spranca, Protected Values, Virology 70 (1997) 1–16. doi: 10.1006/obhd.1997. 2690

  36. [36]

    Dickert, D

    S. Dickert, D. Västfjäll, J. Kleber, P. Slovic, Scope insensitivity: The limits of intuitive valuation of human lives in public policy, Journal of Applied Research in Memory and Cognition 4 (2015) 248–

  37. [37]

    doi: 10.1016/ j.jarmac.2014.09.002

    URL: https://www.sciencedirect.com/science/article/pii/S2211368114000795. doi: 10.1016/ j.jarmac.2014.09.002

  38. [38]

    Brandt (Ed.), Handbook of computational social choice, Cambridge University Press, Cambridge ; New York, 2016

    F. Brandt (Ed.), Handbook of computational social choice, Cambridge University Press, Cambridge ; New York, 2016

  39. [39]

    Arrow, Social Choice and Individual Values, Cowles Foundation Monograph Series, Yale Uni- versity Press, 1970

    K. Arrow, Social Choice and Individual Values, Cowles Foundation Monograph Series, Yale Uni- versity Press, 1970. URL: https://books.google.com/books?id=uebtAAAAMAAJ

  40. [40]

    Conitzer, T

    V. Conitzer, T. Sandholm, Communication complexity of common voting rules, in: Proceedings of the 6th ACM conference on Electronic commerce, ACM, Vancouver BC Canada, 2005, pp. 78–87. URL: https://dl.acm.org/doi/10.1145/1064009.1064018. doi:10.1145/1064009.1064018

  41. [41]

    Conitzer, R

    V. Conitzer, R. Freedman, J. Heitzig, W. H. Holliday, B. M. Jacobs, N. Lambert, M. Mossé, E. Pacuit, S. Russell, H. Schoelkopf, E. Tewolde, W. S. Zwicker, Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback, 2024. URL: http://arxiv.org/abs/2404.10271. doi: 10. 48550/arXiv.2404.10271 , arXiv:2404.10271 [cs]

  42. [42]

    G. E. Gigerenzer, R. E. Hertwig, T. E. Pachur, Heuristics: The foundations of adaptive behavior., Oxford university press, 2011

  43. [43]

    C. R. Sunstein, Social Norms and Social Roles, Columbia Law Review 96 (1996) 903. URL: https: //www.jstor.org/stable/1123430?origin=crossref. doi:10.2307/1123430

  44. [44]

    J. S. Mill, Utilitarianism, in: Seven masterpieces of philosophy, Routledge, 2016, pp. 329–375

  45. [45]

    Kant, Groundwork of the metaphysic of morals, in: Immanuel Kant, Routledge, 2020, pp

    I. Kant, Groundwork of the metaphysic of morals, in: Immanuel Kant, Routledge, 2020, pp. 17–98

  46. [46]

    T. M. Scanlon, What we owe to each other, Belknap Press, 2000

  47. [47]

    Gottlieb, Aristotle: nicomachean ethics, in: Central Works of Philosophy v1, Routledge, 2015, pp

    P. Gottlieb, Aristotle: nicomachean ethics, in: Central Works of Philosophy v1, Routledge, 2015, pp. 46–68

  48. [48]

    Gilligan, In a different voice: Psychological theory and women’s development, Harvard uni- versity press, 1993

    C. Gilligan, In a different voice: Psychological theory and women’s development, Harvard uni- versity press, 1993. A. Formal Definitions and Notation Symbol / Term Name Definition / Role 𝐺⋆ = (𝑉⋆, 𝐸⋆) Moral interaction graph Full graph of morally relevant entities and their local influence relations. 𝑣 ∈ 𝑉 ⋆ Moral entity Node whose state contributes direc...

  49. [49]

    Depth Reversal: Truncated evaluation ( 𝐻 small) favors aggressive intervention, while longer rollouts (𝐻 large) favor targeted intervention

  50. [50]

    Breadth Constraint: Coarse abstraction can eliminate the targeted intervention from the ad- missible action set. B.1. Instantiation of the Moral System Moral Interaction Graph. Let 𝑉 ⋆ = {𝐸1, 𝐸2, 𝐶, 𝑀1, 𝑀2}, representing two extremist nodes, a connector node, and two moderate nodes. The moral interaction graph 𝐺⋆ = (𝑉⋆, 𝐸⋆)is 𝐸⋆ = {{𝐸1, 𝐸2}, {𝐸1, 𝐶}, {𝐸2,...