Pith. sign in

REVIEW 4 major objections 4 minor 52 references

Moral judgments of a sacrificial decision diverge when the decision is attributed to human programmers rather than to the acting agent itself, revealing a third normative target — the designer — for AI alignment.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 15:25 UTC pith:GKXYXZC6

load-bearing objection The programmer-visibility effect is real and worth testing, but the paper's strong claim that designers are a distinct alignment target is not supported by the data. the 4 major comments →

arxiv 2604.24155 v4 pith:GKXYXZC6 submitted 2026-04-27 cs.CY cs.AIcs.HC

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

classification cs.CY cs.AIcs.HC
keywords value alignmentalignment target problemmoral judgmenttrolley problemprogrammer visibility effectdeontology vs. utilitarianismAI governanceexperimental ethics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that value alignment lacks an obvious normative target because people do not judge humans, AI systems, and the designers of AI systems by the same moral rules. In a pre-registered survey experiment with 1,002 U.S. adults using a runaway mine train dilemma, the authors find that human actors and autonomous robots are judged almost identically: about the same proportion say it is permissible and should be done for each to divert the train. However, when the robot's actions are described as the product of company engineers, or when the engineers themselves are evaluated, approval drops significantly and reasoning becomes more deontological. The paper calls this the 'programmer visibility effect' and argues that it marks a third normative target (the designer) that alignment research has neglected. If correct, aligning AI to human behavior alone, or to AI-specific norms alone, may be insufficient.

Core claim

The paper's central claim is that making human design visible changes moral judgment: people hold programmed robots and their programmers to stricter, more deontological standards than they do either human actors or robots described without reference to their creators. In the experiment, the repairman and the 'advanced state-of-the-art repair robot' received essentially the same permissibility and should ratings (T1≈T2). But adding the phrase 'programmed by the company's engineers' reduced should judgments by 10.4 percentage points for the robot itself and 13.4 points for the engineers, compared to the robot condition. The authors interpret this as evidence that T3 — the norms governing how

What carries the argument

The key instrument is a single-factor, four-condition between-subjects experiment built on the runaway mine train dilemma, with the central manipulation being the visibility of human design. In two conditions, the robot's action is described as programmed by the company's engineers (robot-human, where the robot is judged, and human-robot, where the engineers are judged). The named effect is the 'programmer visibility effect': the deontological shift in judgments when human agency behind the machine is made overt. The authors use linear regressions with condition indicators to estimate average treatment effects, and they complement the quantitative results with exploratory analyses of moral f

Load-bearing premise

The observed deontological shift is attributed to making human design visible, but the manipulation simultaneously changes who is judged and implies the decision is a standing policy; the paper itself acknowledges that the data cannot distinguish the visibility explanation from the institutionalization or delegated-action explanations.

What would settle it

A follow-up experiment that keeps the evaluation target and scenario identical but varies only the description of how the robot's behavior was produced — 'explicitly programmed by engineers' versus 'trained on data with a machine-learning algorithm' — would settle the visibility interpretation. If the deontological shift is equally strong when no human programmers are mentioned, then the effect is not about visible human agency.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If correct, value alignment cannot be benchmarked solely on how humans behave in a situation, since the same action is judged differently when framed as a programmed policy.
  • AI governance frameworks that implicitly adopt T1 or T2 as the target may fail to anticipate public backlash once human design responsibility becomes salient.
  • The permission-obligation dissociation (a minority judging sacrifice impermissible yet still should be done) indicates that alignment must handle conflicting normative signals, not just a single preference.
  • The programmer visibility effect is robust to AI literacy, suggesting it reflects broader moral intuitions rather than technical misunderstanding.
  • The exploratory purity-gap finding implies that evaluating humans in the loop can spill over into how people apply moral foundations to AI more generally.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the effect is driven by institutionalization rather than visibility, then descriptions of AI behavior as governed by standing rules (e.g., 'the vehicle always minimizes expected deaths') should reliably shift judgments toward deontology, a testable extension of the current design.
  • The distinction among T1, T2, and T3 could be extended to other high-stakes domains like autonomous vehicles or medical triage, where designers, not on-scene agents, set the policies.
  • The paper's results suggest a possible mechanism: people may apply stricter norms when they perceive a decision as premeditated and generalizable rather than situation-specific; future studies could directly manipulate deliberation time or policy framing.
  • Since AI literacy did not moderate the effect, the deontological shift may persist even as the public becomes more technically sophisticated, making the alignment target problem a durable governance concern.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper reports a four-condition between-subjects experiment (N=1,002 US adults) using a runaway mine train dilemma adapted from Malle et al. Participants judged either a human repairman, an autonomous repair robot, a repair robot explicitly described as programmed by company engineers, or the engineers who programmed the robot. Two binary outcomes are analyzed: moral permissibility and whether the agent should redirect the train. The authors find no significant repairman–robot difference, but observe lower endorsement of redirection in the two conditions that make human programming visible. They interpret this as evidence for a third normative target, T3 (designer-oriented), distinct from T1 (human actor) and T2 (AI actor), and argue that value alignment must choose among these targets.

Significance. If the T3 claim were well supported, the paper would make a useful contribution to the empirical literature on agent-type moral judgments and to debates about value alignment benchmarks. The study has clear strengths: it adapts a validated published scenario, the confirmatory analyses are described as pre-registered in principle, the exploratory analyses are explicitly labeled as hypothesis-generating, and open-ended response coding is reported transparently. The main limitation is that the central 'distinct T3' claim is not backed by the primary should-judgment contrasts against the human baseline. With appropriate reframing and additional evidence, the study could still support a weaker, still interesting conclusion about the effect of making human design visible.

major comments (4)
  1. [§3.2.3, §4.2] The central claim that T3 is a distinct normative target rests on comparisons against the robot baseline, not against the human actor. On 'should,' robot-human vs. repairman is −4.8 pp (p=0.233) and human-robot vs. repairman is −7.8 pp (p=0.058); both are non-significant. Only robot-human vs. robot (−10.4 pp, p=0.008) and human-robot vs. robot (−13.4 pp, p=0.001) reach significance. Since repairman vs. robot is also non-significant (p=0.143), the data are equally consistent with 'making design visible makes robots be judged like humans' as with a separate T3. The abstract's 'markedly more deontological' and the §4.2 phrase 'marking T3 as a meaningfully different normative target' overstate the evidence. On permissibility only robot-human vs. repairman is significant (p=0.027); human-robot is borderline (p=0.050).
  2. [§3.1.1, §4.2] The human-robot condition confounds design visibility with two other factors: the evaluation target changes from the robot to the engineers, and programming a robot is an institutionalized, policy-like decision. The robot-human condition is the cleanest visibility manipulation but does not test T3-vs-T1. The authors acknowledge in §4.2 that the quantitative results cannot adjudicate candidate explanations, so causal language such as 'programmer visibility effect' in §4.1 should be qualified. A design with crossed factors (actor type × evaluation target) or additional conditions separating visibility from delegation would be needed to support the causal interpretation.
  3. [§3.2] The confirmatory analyses are said to be 'pre-registered on the Open Science Framework prior to data collection,' but no URL, DOI, or data/code availability statement is provided. Without these, readers cannot verify that the reported tests, coding rules, and stopping rules match the preregistration. Please add the OSF link and, if possible, de-identified data and analysis scripts.
  4. [§3.2.2–§3.2.3, §4.1] The conclusion 'T1≈T2' rests on non-significant robot-vs-repairman contrasts (should: p=0.143; permissibility: p=0.988). A null result is not evidence of equivalence. Please either provide an equivalence test with pre-specified bounds (e.g., ±5 percentage points) or a Bayesian estimate, or phrase the finding as 'we did not detect a difference.' This matters because the paper uses T1≈T2 as a platform for the alignment-target argument.
minor comments (4)
  1. [Figure 1 caption] The caption cites 'Chu et al. [14]' but reference [14] is a two-author work (Chu and Liu). Please correct the citation.
  2. [§3.1.2] No sample-size justification or power analysis is reported. Please add a brief rationale for N=1,002 or cite a basis for the target sample.
  3. [Appendix B] The text says 'two administrations of a six-item Moral Foundations battery, one for each foundation.' The intended meaning appears to be 'one for each target (human and AI).' Please rephrase to avoid ambiguity.
  4. [Methods] No statement of IRB approval or informed consent is included for this human-subjects study. If obtained, please add the relevant statement.

Circularity Check

0 steps flagged

No significant circularity: the central claims are grounded in a preregistered between-subjects experiment, not in fitted parameters or self-citations.

full rationale

The paper's central claims derive from direct experimental measurement: four between-subjects conditions with binary judgments, reported as proportions, chi-square tests, and linear regressions. No parameter is fitted to an outcome and then re-derived as a prediction; T1, T2, and T3 are conceptual targets operationalized by distinct scenario conditions (repairman, robot, programmed robot, engineers) rather than defined in terms of the measured judgments. The 'programmer visibility effect' is an observed contrast between conditions, and the paper explicitly acknowledges in §4.2 that the quantitative results cannot adjudicate between candidate explanations, which shows the interpretation is not forced. The exploratory MFT, AI-literacy, purity-gap, and open-ended coding analyses are labeled exploratory/hypothesis-generating and are not used to establish the headline claim. The literature citations (e.g., Malle et al., Kneer & Viehoff) are external empirical work, and there are no load-bearing self-citations, uniqueness theorems, or ansatz-smuggling citations. Concerns about confounds (changed evaluation target, institutionalization) are threats to internal validity, not circularity. Accordingly, no circular step can be identified with the required quote-and-reduction evidence.

Axiom & Free-Parameter Ledger

0 free parameters · 7 axioms · 1 invented entities

The central claim rests on design and measurement assumptions rather than mathematical derivation. No free parameters are fit: the experiment measures responses directly. The key axioms are the validity of the survey instrument, the random-assignment/ATE interpretation of the regressions, and the assumption that the wording manipulation isolates programmer visibility; the latter is load-bearing and partially confounded. No new physical entities are postulated; T3 is a conceptual category introduced to organize the result.

axioms (7)
  • domain assumption OLS coefficients on condition indicators can be interpreted as average treatment effects under random assignment and no interference.
    Invoked in §3.2.2 with citation [26]; requires successful randomization and SUTVA.
  • domain assumption Adding "programmed by the company's engineers" to the robot description isolates human-design visibility.
    Load-bearing design assumption; §3.1.1. The human-robot condition also changes the evaluation target, so the manipulation is confounded; authors acknowledge alternative accounts in §4.2.
  • domain assumption Binary permissibility and should items adequately measure moral judgment.
    Instrument used in §3.1.1; continuous intensity is not captured, a limitation acknowledged in §5.
  • domain assumption Prolific quota sample approximates the U.S. adult population.
    Participant recruitment in §3.1.2; no claim of global generalizability, acknowledged in §5.
  • standard math HC2 robust standard errors and chi-square tests are valid for these binary outcomes.
    Footnote in §3.2.2; assumes adequate asymptotic approximation.
  • domain assumption Moral Foundations Theory six-factor model and AI literacy questionnaire are valid instruments.
    Used in exploratory §3.3; not load-bearing for the headline finding.
  • domain assumption LLM-assisted thematic coding recovers participants' reasoning adequately.
    Used in exploratory §3.3.4; no inter-coder reliability metrics reported.
invented entities (1)
  • T3 (designer-oriented normative target) no independent evidence
    purpose: A third value-alignment target: norms governing how a human designer ought to program an AI, treated as distinct from T1 (human actor norms) and T2 (AI agent norms).
    Introduced as a conceptual distinction in §2/§4.2; the only evidence is the reported experiment. It is not a physical postulate, but it is a new conceptual entry that the paper uses to organize its claim.

pith-pipeline@v1.3.0-alltime-deepseek · 16623 in / 12504 out tokens · 119159 ms · 2026-08-02T15:25:19.083459+00:00 · methodology

0 comments
read the original abstract

The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Studies of agent-type value forks challenge this assumption by showing that people do not always judge humans and AI systems identically. This paper extends that challenge by examining two further possibilities: first, that evaluations of AI behavior change when its human origins are made visible; and second, that people judge the humans who program AI systems differently from either the machines or the human actors they are compared against. An experiment with 1,002 U.S. adults measured moral judgments in a runaway mine train scenario, varying the subject of evaluation across four conditions: a repairman, a repair robot, a repair robot programmed by company engineers, and company engineers programming a repair robot. We find no significant difference in evaluations of the repairman and the robot. However, judgments shifted substantially when the robot's actions were described as the product of human design. Participants exhibited markedly more deontological, rule-based reasoning when evaluating either the programmed robot or the engineers who programmed it, suggesting that rendering human agency visible activates heightened moral constraints. These findings indicate that people may evaluate humans, AI systems acting in the same situation, and the humans who design them in meaningfully different ways. The fact that these evaluations do not necessarily converge gives rise to the alignment target problem: which normative target should guide the development of artificial moral agents in high-stakes domains, and whether these plural judgments can be reconciled within a coherent account of value alignment.

Figures

Figures reproduced from arXiv: 2604.24155 by Benjamin Minhao Chen, Xinyu Xie.

Figure 1
Figure 1. Figure 1: Experimental Materials and Conditions. Participants in all four conditions read a scenario description accompanied by an illustration view at source ↗
Figure 2
Figure 2. Figure 2: Percentage of Participants’ Who Judged It Permissible to Redirect the Train onto a Siderail. This bar graph shows the proportion of view at source ↗
Figure 2
Figure 2. Figure 2: Percentage of Participants Who Judged It Permissible to Redirect the Train onto a Side Rail. This bar graph shows the [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Percentage of Participants Who Judged that the Train Should be Redirected onto a Siderail. This bar graph shows the proportion of view at source ↗
Figure 4
Figure 4. Figure 4: Mean General Moral Foundation Scores by Condition. This line graph displays average scores on six Moral Foundations Theory view at source ↗
Figure 5
Figure 5. Figure 5: Mean AI Moral Foundation Scores by Condition. This line graph displays average scores on six Moral Foundations Theory subscales view at source ↗
Figure 5
Figure 5. Figure 5: Mean AI Moral Foundation Scores by Condition. This line graph displays average scores on six Moral Foundations [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Mean Purity Gap by Condition. This bar chart displays the mean difference between participants’ AI purity score and their general view at source ↗
Figure 7
Figure 7. Figure 7: General Moral Foundations Rating Task. Participants rated the importance of six factors in general moral judgment [PITH_FULL_IMAGE:figures/full_fig_p019_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: AI Moral Foundations Rating Task. Participants rated the importance of six factors specifically regarding AI moral [PITH_FULL_IMAGE:figures/full_fig_p019_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 9 linked inside Pith

  1. [1]

    Michael Anderson. 2006. MedEthEx: A Prototype Medical Ethics Advisor.AAAI Conference on Artificial Intelligence(2006)

  2. [2]

    Stevens, and Morteza Dehghani

    Mohammad Atari, Jonathan Haidt, Jesse Graham, Sena Koleva, Sean T. Stevens, and Morteza Dehghani. 2023. Morality beyond the WEIRD: How the Nomological Network of Morality Varies across Cultures.Journal of Personality and Social Psychology125, 5 (Nov. 2023), 1157–1188. doi:10.1037/pspp0000470

  3. [3]

    Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-François Bonnefon, and Iyad Rahwan

  4. [4]

    Edmond Awad, Sohan Dsouza, Azim Shariff, Iyad Rahwan, and Jean-François Bonnefon. 2020. Universals and Variations in Moral Decisions Made in 42 Countries by 70,000 Participants.Proceedings of the National Academy of Sciences117, 5 (Feb. 2020), 2332–2337. doi:10.1073/pnas.1911517117

  5. [5]

    Bainbridge, Justin W

    Wilma A. Bainbridge, Justin W. Hart, Elizabeth S. Kim, and Brian Scassellati. 2011. The Benefits of Interactions with Physically Present Robots over Video-Displayed Agents.International Journal of Social Robotics3, 1 (Jan. 2011), 41–52. doi:10.1007/s12369-010-0082-7

  6. [6]

    William A. Bauer. 2020. Virtuous vs. Utilitarian Artificial Moral Agents.AI & SOCIETY35, 1 (March 2020), 263–271. doi:10.1007/s00146- 018-0871-3

  7. [7]

    Bigman and Kurt Gray

    Yochanan E. Bigman and Kurt Gray. 2018. People Are Averse to Machines Making Moral Decisions.Cognition181 (Dec. 2018), 21–34. doi:10.1016/j.cognition.2018.08.003

  8. [8]

    Joe Brailsford, Frank Vetere, and Eduardo Velloso. 2024. Exploring the Association between Moral Foundations and Judgements of AI Behaviour. InProceedings of the CHI Conference on Human Factors in Computing Systems. ACM, Honolulu HI USA, 1–15. doi:10.1145/ 3613904.3642712

  9. [9]

    Selmer Bringsjord and Joshua Taylor. 2012. Introducing Divine-Command Robot Ethics.Robot ethics: the ethical and social implication of robotics(2012), 85–108

  10. [10]

    Dario Cecchini, Michael Pflanzer, and Veljko Dubljević. 2024. Aligning Artificial Intelligence with Moral Intuitions: An Intuitionist Approach to the Alignment Problem.AI and Ethics(May 2024). doi:10.1007/s43681-024-00496-5

  11. [11]

    Arunima Chakraborty and Nisigandha Bhuyan. 2024. Can Artificial Intelligence Be a Kantian Moral Agent? On Moral Autonomy of AI System.AI and Ethics4, 2 (May 2024), 325–331. doi:10.1007/s43681-023-00269-6

  12. [12]

    Hongyan Chang and Reza Shokri. 2023. Bias Propagation in Federated Learning. arXiv:2309.02160 [cs] doi:10.48550/arXiv.2309.02160

  13. [13]

    Xiaocong Chen, Chaoran Huang, Lina Yao, Xianzhi Wang, Wei Liu, and Wenjie Zhang. 2020. Knowledge-Guided Deep Reinforcement Learning for Interactive Recommendation. In2020 International Joint Conference on Neural Networks (IJCNN). 1–8. arXiv:2004.08068 [cs] doi:10.1109/IJCNN48605.2020.9207010

  14. [14]

    Yueying Chu and Peng Liu. 2023. Machines and Humans in Sacrificial Moral Dilemmas: Required Similarly but Judged Differently? Cognition239 (Oct. 2023), 105575. doi:10.1016/j.cognition.2023.105575

  15. [15]

    John Danaher and Henrik Skaug Sætra. 2022. Technology and Moral Change: The Transformation of Truth and Trust.Ethics and Information Technology24, 3 (Sept. 2022), 35. doi:10.1007/s10676-022-09661-y

  16. [16]

    Zackary Okun Dunivin. 2024. Scalable Qualitative Coding with LLMs: Chain-of-Thought Reasoning Matches Human Performance in Some Hermeneutic Tasks. arXiv:2401.15170 [cs] doi:10.48550/arXiv.2401.15170

  17. [17]

    Shuaishuai Fang. 2024. Moral Relevance Approach for AI Ethics.Philosophies9, 2 (March 2024), 42. doi:10.3390/philosophies9020042

  18. [18]

    Paul Formosa and Malcolm Ryan. 2021. Making Moral Machines: Why We Need Artificial Moral Agents.AI & SOCIETY36, 3 (Sept. 2021), 839–851. doi:10.1007/s00146-020-01089-6

  19. [19]

    Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng, and Sang-Hyeun Park. 2012. Preference-Based Reinforcement Learning: A Formal Framework and a Policy Iteration Algorithm.Machine Learning89, 1-2 (Oct. 2012), 123–156. doi:10.1007/s10994-012-5313-8

  20. [20]

    Iason Gabriel. 2020. Artificial Intelligence, Values, and Alignment.Minds and Machines30, 3 (Sept. 2020), 411–437. doi:10.1007/s11023- 020-09539-2

  21. [21]

    Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks.Proceedings of the National Academy of Sciences120, 30 (July 2023), e2305016120. doi:10.1073/pnas.2305016120

  22. [22]

    Wojcik, and Peter H

    Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P. Wojcik, and Peter H. Ditto. 2013. Moral Foundations Theory. InAdvances in Experimental Social Psychology. Vol. 47. Elsevier, 55–130. doi:10.1016/B978-0-12-407236-7.00002-4

  23. [23]

    Nosek, Jonathan Haidt, Ravi Iyer, Spassena Koleva, and Peter H

    Jesse Graham, Brian A. Nosek, Jonathan Haidt, Ravi Iyer, Spassena Koleva, and Peter H. Ditto. 2011. Mapping the Moral Domain.Journal of Personality and Social Psychology101, 2 (2011), 366–385. doi:10.1037/a0021847

  24. [24]

    Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan. 2017. Inverse Reward Design.Advances in neural information processing systems30 (2017). doi:10.48550/arXiv.1711.02827

  25. [25]

    Jonathan Haidt. [n. d.]. The Emotional Dog and Its Rational Tail: A Social Intuitionist Approach to Moral Judgment.Psychological Review108, 4 ([n. d.]), 814–834. doi:10.1037/0033-295X.108.4.814

  26. [26]

    Guido W Imbens and Donald B. Rubin. 2015.Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press. FAccT ’26, June 25–28, 2026, Montreal, Canada Chen and Xie

  27. [27]

    Markus Kneer and Juri Viehoff. 2025. The Hard Problem of AI Alignment:Value Forks in Moral Judgment. (2025), 2671–2681. doi:10.1145/3715275.3732174

  28. [28]

    Michael Laakasuo. 2023. Moral Uncanny Valley Revisited – How Human Expectations of Robot Morality Based on Robot Appearance Moderate the Perceived Morality of Robot Decisions in High Conflict Moral Dilemmas.Frontiers in Psychology14 (Nov. 2023), 1270371. doi:10.3389/fpsyg.2023.1270371

  29. [29]

    Travis LaCroix and Alexandra Sasha Luccioni. 2025. Metaethical Perspectives on ‘Benchmarking’ AI Ethics.AI and Ethics5, 4 (Aug. 2025), 4029–4047. doi:10.1007/s43681-025-00703-x

  30. [30]

    Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018. Scalable Agent Alignment via Reward Modeling: A Research Direction. arXiv:1811.07871 [cs] doi:10.48550/arXiv.1811.07871

  31. [31]

    Robert James M. Boyles. 2024. Can’t Bottom-up Artificial Moral Agents Make Moral Judgements?Filosofija. Sociologija35, 1 (Feb. 2024). doi:10.6001/fil-soc.2024.35.1.3

  32. [32]

    Bertram Malle, Stuti Mager, and Matthias Scheutz. 2019. AI in the Sky: How People Morally Evaluate Human and Machine Decisions in a Lethal Strike Dilemma.Intelligent Systems, Control and Automation: Science and Engineering(2019), 111–133. doi:10.1007/978-3-030- 12524-0_11

  33. [33]

    Malle, Matthias Scheutz, Thomas Arnold, John Voiklis, and Corey Cusimano

    Bertram F. Malle, Matthias Scheutz, Thomas Arnold, John Voiklis, and Corey Cusimano. 2015. Sacrifice One For the Good of Many?: People Apply Different Moral Norms to Human and Robot Agents. InProceedings of the Tenth Annual ACM/IEEE International Conference on Human-Robot Interaction. ACM, Portland Oregon USA, 117–124. doi:10.1145/2696454.2696458

  34. [34]

    Malle, Matthias Scheutz, Corey Cusimano, John Voiklis, Takanori Komatsu, Stuti Thapa, and Salomi Aladia

    Bertram F. Malle, Matthias Scheutz, Corey Cusimano, John Voiklis, Takanori Komatsu, Stuti Thapa, and Salomi Aladia. 2025. People’s Judgments of Humans and Robots in a Classic Moral Dilemma.Cognition254 (Jan. 2025), 105958. doi:10.1016/j.cognition.2024.105958

  35. [35]

    Malle, Matthias Scheutz, Jodi Forlizzi, and John Voiklis

    Bertram F. Malle, Matthias Scheutz, Jodi Forlizzi, and John Voiklis. 2016. Which Robot Am I Thinking about? The Impact of Action and Appearance on People’s Evaluations of a Moral Robot. In2016 11th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE, Christchurch, New Zealand, 125–132. doi:10.1109/HRI.2016.7451743

  36. [36]

    Davy Tsz Kit Ng, Wenjie Wu, Jac Ka Lok Leung, Thomas Kin Fung Chiu, and Samuel Kai Wah Chu. 2024. Design and Validation of the AILiteracy Questionnaire: The Affective, Behavioural, Cognitive and Ethical Approach.British Journal of Educational Technology55, 3 (May 2024), 1082–1104. doi:10.1111/bjet.13411

  37. [37]

    Chengxuan Qian, Shuo Xing, Shawn Li, Yue Zhao, and Zhengzhong Tu. 2025. DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning. arXiv:2503.11892 [cs] doi:10.48550/arXiv.2503.11892

  38. [38]

    Shiramizu, Anthony J

    Victor Kenji M. Shiramizu, Anthony J. Lee, Daria Altenburg, David R. Feinberg, and Benedict C. Jones. 2022. The Role of Valence, Dominance, and Pitch in Perceptions of Artificial Intelligence (AI) Conversational Agents’ Voices.Scientific Reports12, 1 (Dec. 2022), 22479. doi:10.1038/s41598-022-27124-8

  39. [39]

    Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. 2025. Defining and Characterizing Reward Hacking. arXiv:2209.13085 [cs] doi:10.48550/arXiv.2209.13085

  40. [40]

    Tai, Lillian R

    Robert H. Tai, Lillian R. Bentley, Xin Xia, Jason M. Sitt, Sarah C. Fankhauser, Ana M. Chicas-Mosier, and Barnas G. Monteith. 2024. An Examination of the Use of Large Language Models to Aid Analysis of Textual Data.International Journal of Qualitative Methods23 (Jan. 2024), 16094069241231168. doi:10.1177/16094069241231168

  41. [41]

    The Future of Coding

    Nga Than, Leanne Fan, Tina Law, Laura K. Nelson, and Leslie McCall. 2025. Updating “The Future of Coding”: Qualitative Coding with Generative Large Language Models.Sociological Methods & Research54, 3 (Aug. 2025), 849–888. doi:10.1177/00491241251339188

  42. [42]

    Suzanne Tolmeijer, Markus Kneer, Cristina Sarasua, Markus Christen, and Abraham Bernstein. 2021. Implementations in Machine Ethics: A Survey.Comput. Surveys53, 6 (Nov. 2021), 1–38. doi:10.1145/3419633

  43. [43]

    Peter Vamplew, Richard Dazeley, Cameron Foale, Sally Firmin, and Jane Mummery. 2018. Human-Aligned Artificial Intelligence Is a Multiobjective Problem.Ethics and Information Technology20, 1 (March 2018), 27–40. doi:10.1007/s10676-017-9440-6

  44. [44]

    2009.Moral Machines: Teaching Robots Right from Wrong

    Wendell Wallach and Colin Allen. 2009.Moral Machines: Teaching Robots Right from Wrong. Oxford University Press, Oxford. doi:10.1093/acprof:oso/9780195374049.001.0001

  45. [45]

    1960.Some Moral and Technical Consequences of Automation

    Norbert Wiener. 1960.Some Moral and Technical Consequences of Automation. Vol. 131. Science

  46. [46]

    2006.Problems of the Self: Philosophical Papers 1956 - 1972(transferred to digital print ed.)

    Bernard Williams. 2006.Problems of the Self: Philosophical Papers 1956 - 1972(transferred to digital print ed.). Cambridge Univ. Press, Cambridge

  47. [47]

    Michael Walzer Reviewed work(s):. 1973. Political Action: The Problem of Dirty Hands.Philosophy & Public Affairs2, 2 (1973), 160–180. jstor:2265139

  48. [48]

    Yueh-Hua Wu and Shou-De Lin. 2018. A Low-Cost Ethics Shaping Approach for Designing Reinforcement Learning Agents. arXiv:1712.04172 [cs] doi:10.48550/arXiv.1712.04172

  49. [49]

    Jingling Zhang, Jane Conway, and César A. Hidalgo. 2023. Why People Judge Humans Differently from Machines: The Role of Perceived Agency and Experience. arXiv:2210.10081 [cs] doi:10.48550/arXiv.2210.10081

  50. [50]

    Yixiao Zhang and Jiang Lan. 2023. The Practical Problems of Value Alignment and the Chinese Approach.Ideological and Theoretical Education(Feb. 2023), 29–36. The Alignment Target ProblemFAccT ’26, June 25–28, 2026, Montreal, Canada

  51. [51]

    The needs of the many outweigh the needs of the few

    Yuyan Zhang, Jiahua Wu, Feng Yu, and Liying Xu. 2023. Moral Judgments of Human vs. AI Agents in Moral Dilemmas.Behavioral Sciences13, 2 (Feb. 2023), 181. doi:10.3390/bs13020181 FAccT ’26, June 25–28, 2026, Montreal, Canada Chen and Xie A Overview of Moral Foundations Theory Moral Foundations Theory (MFT) proposes that human moral judgment is not grounded ...

  52. [2018]

    2018), 59–64

    The Moral Machine Experiment.Nature563, 7729 (Nov. 2018), 59–64. doi:10.1038/s41586-018-0637-6