REVIEW 4 major objections 4 minor 52 references
Moral judgments of a sacrificial decision diverge when the decision is attributed to human programmers rather than to the acting agent itself, revealing a third normative target — the designer — for AI alignment.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 15:25 UTC pith:GKXYXZC6
load-bearing objection The programmer-visibility effect is real and worth testing, but the paper's strong claim that designers are a distinct alignment target is not supported by the data. the 4 major comments →
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that making human design visible changes moral judgment: people hold programmed robots and their programmers to stricter, more deontological standards than they do either human actors or robots described without reference to their creators. In the experiment, the repairman and the 'advanced state-of-the-art repair robot' received essentially the same permissibility and should ratings (T1≈T2). But adding the phrase 'programmed by the company's engineers' reduced should judgments by 10.4 percentage points for the robot itself and 13.4 points for the engineers, compared to the robot condition. The authors interpret this as evidence that T3 — the norms governing how
What carries the argument
The key instrument is a single-factor, four-condition between-subjects experiment built on the runaway mine train dilemma, with the central manipulation being the visibility of human design. In two conditions, the robot's action is described as programmed by the company's engineers (robot-human, where the robot is judged, and human-robot, where the engineers are judged). The named effect is the 'programmer visibility effect': the deontological shift in judgments when human agency behind the machine is made overt. The authors use linear regressions with condition indicators to estimate average treatment effects, and they complement the quantitative results with exploratory analyses of moral f
Load-bearing premise
The observed deontological shift is attributed to making human design visible, but the manipulation simultaneously changes who is judged and implies the decision is a standing policy; the paper itself acknowledges that the data cannot distinguish the visibility explanation from the institutionalization or delegated-action explanations.
What would settle it
A follow-up experiment that keeps the evaluation target and scenario identical but varies only the description of how the robot's behavior was produced — 'explicitly programmed by engineers' versus 'trained on data with a machine-learning algorithm' — would settle the visibility interpretation. If the deontological shift is equally strong when no human programmers are mentioned, then the effect is not about visible human agency.
If this is right
- If correct, value alignment cannot be benchmarked solely on how humans behave in a situation, since the same action is judged differently when framed as a programmed policy.
- AI governance frameworks that implicitly adopt T1 or T2 as the target may fail to anticipate public backlash once human design responsibility becomes salient.
- The permission-obligation dissociation (a minority judging sacrifice impermissible yet still should be done) indicates that alignment must handle conflicting normative signals, not just a single preference.
- The programmer visibility effect is robust to AI literacy, suggesting it reflects broader moral intuitions rather than technical misunderstanding.
- The exploratory purity-gap finding implies that evaluating humans in the loop can spill over into how people apply moral foundations to AI more generally.
Where Pith is reading between the lines
- If the effect is driven by institutionalization rather than visibility, then descriptions of AI behavior as governed by standing rules (e.g., 'the vehicle always minimizes expected deaths') should reliably shift judgments toward deontology, a testable extension of the current design.
- The distinction among T1, T2, and T3 could be extended to other high-stakes domains like autonomous vehicles or medical triage, where designers, not on-scene agents, set the policies.
- The paper's results suggest a possible mechanism: people may apply stricter norms when they perceive a decision as premeditated and generalizable rather than situation-specific; future studies could directly manipulate deliberation time or policy framing.
- Since AI literacy did not moderate the effect, the deontological shift may persist even as the public becomes more technically sophisticated, making the alignment target problem a durable governance concern.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a four-condition between-subjects experiment (N=1,002 US adults) using a runaway mine train dilemma adapted from Malle et al. Participants judged either a human repairman, an autonomous repair robot, a repair robot explicitly described as programmed by company engineers, or the engineers who programmed the robot. Two binary outcomes are analyzed: moral permissibility and whether the agent should redirect the train. The authors find no significant repairman–robot difference, but observe lower endorsement of redirection in the two conditions that make human programming visible. They interpret this as evidence for a third normative target, T3 (designer-oriented), distinct from T1 (human actor) and T2 (AI actor), and argue that value alignment must choose among these targets.
Significance. If the T3 claim were well supported, the paper would make a useful contribution to the empirical literature on agent-type moral judgments and to debates about value alignment benchmarks. The study has clear strengths: it adapts a validated published scenario, the confirmatory analyses are described as pre-registered in principle, the exploratory analyses are explicitly labeled as hypothesis-generating, and open-ended response coding is reported transparently. The main limitation is that the central 'distinct T3' claim is not backed by the primary should-judgment contrasts against the human baseline. With appropriate reframing and additional evidence, the study could still support a weaker, still interesting conclusion about the effect of making human design visible.
major comments (4)
- [§3.2.3, §4.2] The central claim that T3 is a distinct normative target rests on comparisons against the robot baseline, not against the human actor. On 'should,' robot-human vs. repairman is −4.8 pp (p=0.233) and human-robot vs. repairman is −7.8 pp (p=0.058); both are non-significant. Only robot-human vs. robot (−10.4 pp, p=0.008) and human-robot vs. robot (−13.4 pp, p=0.001) reach significance. Since repairman vs. robot is also non-significant (p=0.143), the data are equally consistent with 'making design visible makes robots be judged like humans' as with a separate T3. The abstract's 'markedly more deontological' and the §4.2 phrase 'marking T3 as a meaningfully different normative target' overstate the evidence. On permissibility only robot-human vs. repairman is significant (p=0.027); human-robot is borderline (p=0.050).
- [§3.1.1, §4.2] The human-robot condition confounds design visibility with two other factors: the evaluation target changes from the robot to the engineers, and programming a robot is an institutionalized, policy-like decision. The robot-human condition is the cleanest visibility manipulation but does not test T3-vs-T1. The authors acknowledge in §4.2 that the quantitative results cannot adjudicate candidate explanations, so causal language such as 'programmer visibility effect' in §4.1 should be qualified. A design with crossed factors (actor type × evaluation target) or additional conditions separating visibility from delegation would be needed to support the causal interpretation.
- [§3.2] The confirmatory analyses are said to be 'pre-registered on the Open Science Framework prior to data collection,' but no URL, DOI, or data/code availability statement is provided. Without these, readers cannot verify that the reported tests, coding rules, and stopping rules match the preregistration. Please add the OSF link and, if possible, de-identified data and analysis scripts.
- [§3.2.2–§3.2.3, §4.1] The conclusion 'T1≈T2' rests on non-significant robot-vs-repairman contrasts (should: p=0.143; permissibility: p=0.988). A null result is not evidence of equivalence. Please either provide an equivalence test with pre-specified bounds (e.g., ±5 percentage points) or a Bayesian estimate, or phrase the finding as 'we did not detect a difference.' This matters because the paper uses T1≈T2 as a platform for the alignment-target argument.
minor comments (4)
- [Figure 1 caption] The caption cites 'Chu et al. [14]' but reference [14] is a two-author work (Chu and Liu). Please correct the citation.
- [§3.1.2] No sample-size justification or power analysis is reported. Please add a brief rationale for N=1,002 or cite a basis for the target sample.
- [Appendix B] The text says 'two administrations of a six-item Moral Foundations battery, one for each foundation.' The intended meaning appears to be 'one for each target (human and AI).' Please rephrase to avoid ambiguity.
- [Methods] No statement of IRB approval or informed consent is included for this human-subjects study. If obtained, please add the relevant statement.
Circularity Check
No significant circularity: the central claims are grounded in a preregistered between-subjects experiment, not in fitted parameters or self-citations.
full rationale
The paper's central claims derive from direct experimental measurement: four between-subjects conditions with binary judgments, reported as proportions, chi-square tests, and linear regressions. No parameter is fitted to an outcome and then re-derived as a prediction; T1, T2, and T3 are conceptual targets operationalized by distinct scenario conditions (repairman, robot, programmed robot, engineers) rather than defined in terms of the measured judgments. The 'programmer visibility effect' is an observed contrast between conditions, and the paper explicitly acknowledges in §4.2 that the quantitative results cannot adjudicate between candidate explanations, which shows the interpretation is not forced. The exploratory MFT, AI-literacy, purity-gap, and open-ended coding analyses are labeled exploratory/hypothesis-generating and are not used to establish the headline claim. The literature citations (e.g., Malle et al., Kneer & Viehoff) are external empirical work, and there are no load-bearing self-citations, uniqueness theorems, or ansatz-smuggling citations. Concerns about confounds (changed evaluation target, institutionalization) are threats to internal validity, not circularity. Accordingly, no circular step can be identified with the required quote-and-reduction evidence.
Axiom & Free-Parameter Ledger
axioms (7)
- domain assumption OLS coefficients on condition indicators can be interpreted as average treatment effects under random assignment and no interference.
- domain assumption Adding "programmed by the company's engineers" to the robot description isolates human-design visibility.
- domain assumption Binary permissibility and should items adequately measure moral judgment.
- domain assumption Prolific quota sample approximates the U.S. adult population.
- standard math HC2 robust standard errors and chi-square tests are valid for these binary outcomes.
- domain assumption Moral Foundations Theory six-factor model and AI literacy questionnaire are valid instruments.
- domain assumption LLM-assisted thematic coding recovers participants' reasoning adequately.
invented entities (1)
-
T3 (designer-oriented normative target)
no independent evidence
read the original abstract
The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Studies of agent-type value forks challenge this assumption by showing that people do not always judge humans and AI systems identically. This paper extends that challenge by examining two further possibilities: first, that evaluations of AI behavior change when its human origins are made visible; and second, that people judge the humans who program AI systems differently from either the machines or the human actors they are compared against. An experiment with 1,002 U.S. adults measured moral judgments in a runaway mine train scenario, varying the subject of evaluation across four conditions: a repairman, a repair robot, a repair robot programmed by company engineers, and company engineers programming a repair robot. We find no significant difference in evaluations of the repairman and the robot. However, judgments shifted substantially when the robot's actions were described as the product of human design. Participants exhibited markedly more deontological, rule-based reasoning when evaluating either the programmed robot or the engineers who programmed it, suggesting that rendering human agency visible activates heightened moral constraints. These findings indicate that people may evaluate humans, AI systems acting in the same situation, and the humans who design them in meaningfully different ways. The fact that these evaluations do not necessarily converge gives rise to the alignment target problem: which normative target should guide the development of artificial moral agents in high-stakes domains, and whether these plural judgments can be reconciled within a coherent account of value alignment.
Figures
Reference graph
Works this paper leans on
-
[1]
Michael Anderson. 2006. MedEthEx: A Prototype Medical Ethics Advisor.AAAI Conference on Artificial Intelligence(2006)
2006
-
[2]
Mohammad Atari, Jonathan Haidt, Jesse Graham, Sena Koleva, Sean T. Stevens, and Morteza Dehghani. 2023. Morality beyond the WEIRD: How the Nomological Network of Morality Varies across Cultures.Journal of Personality and Social Psychology125, 5 (Nov. 2023), 1157–1188. doi:10.1037/pspp0000470
-
[3]
Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-François Bonnefon, and Iyad Rahwan
-
[4]
Edmond Awad, Sohan Dsouza, Azim Shariff, Iyad Rahwan, and Jean-François Bonnefon. 2020. Universals and Variations in Moral Decisions Made in 42 Countries by 70,000 Participants.Proceedings of the National Academy of Sciences117, 5 (Feb. 2020), 2332–2337. doi:10.1073/pnas.1911517117
-
[5]
Wilma A. Bainbridge, Justin W. Hart, Elizabeth S. Kim, and Brian Scassellati. 2011. The Benefits of Interactions with Physically Present Robots over Video-Displayed Agents.International Journal of Social Robotics3, 1 (Jan. 2011), 41–52. doi:10.1007/s12369-010-0082-7
-
[6]
William A. Bauer. 2020. Virtuous vs. Utilitarian Artificial Moral Agents.AI & SOCIETY35, 1 (March 2020), 263–271. doi:10.1007/s00146- 018-0871-3
doi:10.1007/s00146- 2020
-
[7]
Yochanan E. Bigman and Kurt Gray. 2018. People Are Averse to Machines Making Moral Decisions.Cognition181 (Dec. 2018), 21–34. doi:10.1016/j.cognition.2018.08.003
-
[8]
Joe Brailsford, Frank Vetere, and Eduardo Velloso. 2024. Exploring the Association between Moral Foundations and Judgements of AI Behaviour. InProceedings of the CHI Conference on Human Factors in Computing Systems. ACM, Honolulu HI USA, 1–15. doi:10.1145/ 3613904.3642712
arXiv 2024
-
[9]
Selmer Bringsjord and Joshua Taylor. 2012. Introducing Divine-Command Robot Ethics.Robot ethics: the ethical and social implication of robotics(2012), 85–108
2012
-
[10]
Dario Cecchini, Michael Pflanzer, and Veljko Dubljević. 2024. Aligning Artificial Intelligence with Moral Intuitions: An Intuitionist Approach to the Alignment Problem.AI and Ethics(May 2024). doi:10.1007/s43681-024-00496-5
-
[11]
Arunima Chakraborty and Nisigandha Bhuyan. 2024. Can Artificial Intelligence Be a Kantian Moral Agent? On Moral Autonomy of AI System.AI and Ethics4, 2 (May 2024), 325–331. doi:10.1007/s43681-023-00269-6
-
[12]
Hongyan Chang and Reza Shokri. 2023. Bias Propagation in Federated Learning. arXiv:2309.02160 [cs] doi:10.48550/arXiv.2309.02160
-
[13]
Xiaocong Chen, Chaoran Huang, Lina Yao, Xianzhi Wang, Wei Liu, and Wenjie Zhang. 2020. Knowledge-Guided Deep Reinforcement Learning for Interactive Recommendation. In2020 International Joint Conference on Neural Networks (IJCNN). 1–8. arXiv:2004.08068 [cs] doi:10.1109/IJCNN48605.2020.9207010
Pith/arXiv arXiv 2020
-
[14]
Yueying Chu and Peng Liu. 2023. Machines and Humans in Sacrificial Moral Dilemmas: Required Similarly but Judged Differently? Cognition239 (Oct. 2023), 105575. doi:10.1016/j.cognition.2023.105575
arXiv 2023
-
[15]
John Danaher and Henrik Skaug Sætra. 2022. Technology and Moral Change: The Transformation of Truth and Trust.Ethics and Information Technology24, 3 (Sept. 2022), 35. doi:10.1007/s10676-022-09661-y
-
[16]
Zackary Okun Dunivin. 2024. Scalable Qualitative Coding with LLMs: Chain-of-Thought Reasoning Matches Human Performance in Some Hermeneutic Tasks. arXiv:2401.15170 [cs] doi:10.48550/arXiv.2401.15170
-
[17]
Shuaishuai Fang. 2024. Moral Relevance Approach for AI Ethics.Philosophies9, 2 (March 2024), 42. doi:10.3390/philosophies9020042
-
[18]
Paul Formosa and Malcolm Ryan. 2021. Making Moral Machines: Why We Need Artificial Moral Agents.AI & SOCIETY36, 3 (Sept. 2021), 839–851. doi:10.1007/s00146-020-01089-6
-
[19]
Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng, and Sang-Hyeun Park. 2012. Preference-Based Reinforcement Learning: A Formal Framework and a Policy Iteration Algorithm.Machine Learning89, 1-2 (Oct. 2012), 123–156. doi:10.1007/s10994-012-5313-8
-
[20]
Iason Gabriel. 2020. Artificial Intelligence, Values, and Alignment.Minds and Machines30, 3 (Sept. 2020), 411–437. doi:10.1007/s11023- 020-09539-2
doi:10.1007/s11023- 2020
-
[21]
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks.Proceedings of the National Academy of Sciences120, 30 (July 2023), e2305016120. doi:10.1073/pnas.2305016120
-
[22]
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P. Wojcik, and Peter H. Ditto. 2013. Moral Foundations Theory. InAdvances in Experimental Social Psychology. Vol. 47. Elsevier, 55–130. doi:10.1016/B978-0-12-407236-7.00002-4
-
[23]
Nosek, Jonathan Haidt, Ravi Iyer, Spassena Koleva, and Peter H
Jesse Graham, Brian A. Nosek, Jonathan Haidt, Ravi Iyer, Spassena Koleva, and Peter H. Ditto. 2011. Mapping the Moral Domain.Journal of Personality and Social Psychology101, 2 (2011), 366–385. doi:10.1037/a0021847
doi:10.1037/a0021847 2011
-
[24]
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan. 2017. Inverse Reward Design.Advances in neural information processing systems30 (2017). doi:10.48550/arXiv.1711.02827
-
[25]
Jonathan Haidt. [n. d.]. The Emotional Dog and Its Rational Tail: A Social Intuitionist Approach to Moral Judgment.Psychological Review108, 4 ([n. d.]), 814–834. doi:10.1037/0033-295X.108.4.814
-
[26]
Guido W Imbens and Donald B. Rubin. 2015.Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press. FAccT ’26, June 25–28, 2026, Montreal, Canada Chen and Xie
2015
-
[27]
Markus Kneer and Juri Viehoff. 2025. The Hard Problem of AI Alignment:Value Forks in Moral Judgment. (2025), 2671–2681. doi:10.1145/3715275.3732174
arXiv 2025
-
[28]
Michael Laakasuo. 2023. Moral Uncanny Valley Revisited – How Human Expectations of Robot Morality Based on Robot Appearance Moderate the Perceived Morality of Robot Decisions in High Conflict Moral Dilemmas.Frontiers in Psychology14 (Nov. 2023), 1270371. doi:10.3389/fpsyg.2023.1270371
arXiv 2023
-
[29]
Travis LaCroix and Alexandra Sasha Luccioni. 2025. Metaethical Perspectives on ‘Benchmarking’ AI Ethics.AI and Ethics5, 4 (Aug. 2025), 4029–4047. doi:10.1007/s43681-025-00703-x
-
[30]
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018. Scalable Agent Alignment via Reward Modeling: A Research Direction. arXiv:1811.07871 [cs] doi:10.48550/arXiv.1811.07871
-
[31]
Robert James M. Boyles. 2024. Can’t Bottom-up Artificial Moral Agents Make Moral Judgements?Filosofija. Sociologija35, 1 (Feb. 2024). doi:10.6001/fil-soc.2024.35.1.3
-
[32]
Bertram Malle, Stuti Mager, and Matthias Scheutz. 2019. AI in the Sky: How People Morally Evaluate Human and Machine Decisions in a Lethal Strike Dilemma.Intelligent Systems, Control and Automation: Science and Engineering(2019), 111–133. doi:10.1007/978-3-030- 12524-0_11
-
[33]
Malle, Matthias Scheutz, Thomas Arnold, John Voiklis, and Corey Cusimano
Bertram F. Malle, Matthias Scheutz, Thomas Arnold, John Voiklis, and Corey Cusimano. 2015. Sacrifice One For the Good of Many?: People Apply Different Moral Norms to Human and Robot Agents. InProceedings of the Tenth Annual ACM/IEEE International Conference on Human-Robot Interaction. ACM, Portland Oregon USA, 117–124. doi:10.1145/2696454.2696458
arXiv 2015
-
[34]
Bertram F. Malle, Matthias Scheutz, Corey Cusimano, John Voiklis, Takanori Komatsu, Stuti Thapa, and Salomi Aladia. 2025. People’s Judgments of Humans and Robots in a Classic Moral Dilemma.Cognition254 (Jan. 2025), 105958. doi:10.1016/j.cognition.2024.105958
arXiv 2025
-
[35]
Malle, Matthias Scheutz, Jodi Forlizzi, and John Voiklis
Bertram F. Malle, Matthias Scheutz, Jodi Forlizzi, and John Voiklis. 2016. Which Robot Am I Thinking about? The Impact of Action and Appearance on People’s Evaluations of a Moral Robot. In2016 11th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE, Christchurch, New Zealand, 125–132. doi:10.1109/HRI.2016.7451743
arXiv 2016
-
[36]
Davy Tsz Kit Ng, Wenjie Wu, Jac Ka Lok Leung, Thomas Kin Fung Chiu, and Samuel Kai Wah Chu. 2024. Design and Validation of the AILiteracy Questionnaire: The Affective, Behavioural, Cognitive and Ethical Approach.British Journal of Educational Technology55, 3 (May 2024), 1082–1104. doi:10.1111/bjet.13411
-
[37]
Chengxuan Qian, Shuo Xing, Shawn Li, Yue Zhao, and Zhengzhong Tu. 2025. DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning. arXiv:2503.11892 [cs] doi:10.48550/arXiv.2503.11892
-
[38]
Victor Kenji M. Shiramizu, Anthony J. Lee, Daria Altenburg, David R. Feinberg, and Benedict C. Jones. 2022. The Role of Valence, Dominance, and Pitch in Perceptions of Artificial Intelligence (AI) Conversational Agents’ Voices.Scientific Reports12, 1 (Dec. 2022), 22479. doi:10.1038/s41598-022-27124-8
-
[39]
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. 2025. Defining and Characterizing Reward Hacking. arXiv:2209.13085 [cs] doi:10.48550/arXiv.2209.13085
-
[40]
Robert H. Tai, Lillian R. Bentley, Xin Xia, Jason M. Sitt, Sarah C. Fankhauser, Ana M. Chicas-Mosier, and Barnas G. Monteith. 2024. An Examination of the Use of Large Language Models to Aid Analysis of Textual Data.International Journal of Qualitative Methods23 (Jan. 2024), 16094069241231168. doi:10.1177/16094069241231168
-
[41]
Nga Than, Leanne Fan, Tina Law, Laura K. Nelson, and Leslie McCall. 2025. Updating “The Future of Coding”: Qualitative Coding with Generative Large Language Models.Sociological Methods & Research54, 3 (Aug. 2025), 849–888. doi:10.1177/00491241251339188
-
[42]
Suzanne Tolmeijer, Markus Kneer, Cristina Sarasua, Markus Christen, and Abraham Bernstein. 2021. Implementations in Machine Ethics: A Survey.Comput. Surveys53, 6 (Nov. 2021), 1–38. doi:10.1145/3419633
doi:10.1145/3419633 2021
-
[43]
Peter Vamplew, Richard Dazeley, Cameron Foale, Sally Firmin, and Jane Mummery. 2018. Human-Aligned Artificial Intelligence Is a Multiobjective Problem.Ethics and Information Technology20, 1 (March 2018), 27–40. doi:10.1007/s10676-017-9440-6
-
[44]
2009.Moral Machines: Teaching Robots Right from Wrong
Wendell Wallach and Colin Allen. 2009.Moral Machines: Teaching Robots Right from Wrong. Oxford University Press, Oxford. doi:10.1093/acprof:oso/9780195374049.001.0001
arXiv 2009
-
[45]
1960.Some Moral and Technical Consequences of Automation
Norbert Wiener. 1960.Some Moral and Technical Consequences of Automation. Vol. 131. Science
1960
-
[46]
2006.Problems of the Self: Philosophical Papers 1956 - 1972(transferred to digital print ed.)
Bernard Williams. 2006.Problems of the Self: Philosophical Papers 1956 - 1972(transferred to digital print ed.). Cambridge Univ. Press, Cambridge
2006
-
[47]
Michael Walzer Reviewed work(s):. 1973. Political Action: The Problem of Dirty Hands.Philosophy & Public Affairs2, 2 (1973), 160–180. jstor:2265139
1973
-
[48]
Yueh-Hua Wu and Shou-De Lin. 2018. A Low-Cost Ethics Shaping Approach for Designing Reinforcement Learning Agents. arXiv:1712.04172 [cs] doi:10.48550/arXiv.1712.04172
-
[49]
Jingling Zhang, Jane Conway, and César A. Hidalgo. 2023. Why People Judge Humans Differently from Machines: The Role of Perceived Agency and Experience. arXiv:2210.10081 [cs] doi:10.48550/arXiv.2210.10081
-
[50]
Yixiao Zhang and Jiang Lan. 2023. The Practical Problems of Value Alignment and the Chinese Approach.Ideological and Theoretical Education(Feb. 2023), 29–36. The Alignment Target ProblemFAccT ’26, June 25–28, 2026, Montreal, Canada
2023
-
[51]
The needs of the many outweigh the needs of the few
Yuyan Zhang, Jiahua Wu, Feng Yu, and Liying Xu. 2023. Moral Judgments of Human vs. AI Agents in Moral Dilemmas.Behavioral Sciences13, 2 (Feb. 2023), 181. doi:10.3390/bs13020181 FAccT ’26, June 25–28, 2026, Montreal, Canada Chen and Xie A Overview of Moral Foundations Theory Moral Foundations Theory (MFT) proposes that human moral judgment is not grounded ...
-
[2018]
The Moral Machine Experiment.Nature563, 7729 (Nov. 2018), 59–64. doi:10.1038/s41586-018-0637-6
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.