REVIEW 3 major objections 5 minor 7 references
The Emotional Alignment Design Policy
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proposes that artificial entities should be designed to elicit emotional reactions that match their real capacities and moral status.
desk verdict A clear, honest normative proposal that names a new design obligation; the empirical controllability gap is real but not fatal, and the paper deserves a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the distinction between two axes of alignment: degree, meaning avoiding overshooting and undershooting, and type, meaning avoiding wrong-target reactions, together with the idea of fittingness from emotion theory. The paper uses this distinction to generate concrete design prescriptions: align emotion with belief, respect user autonomy, reflect epistemic uncertainty, correct for asymmetrical risk, handle creation and destruction, and mitigate human bias. The key mechanism is that design features such as faces, voices, labels, and behavior are treated as controllable levers that can set a user's felt moral concern to match the system's actual moral standing.
What would settle it
Run a between-subjects experiment where the same underlying AI capacity is presented in bland, cute, and balanced interfaces; if users' emotional engagement and willingness to harm or help the system do not track the actual capacity, or if a calibrated interface consistently triggers uncanny-valley rejection, then the policy lacks the design control it assumes.
Extended reading notes
Core claim
The authors aim to establish a normative design principle: designers of artificial entities should treat emotional alignment as a moral requirement, not merely a user-experience preference. A well-designed AI should make users feel about it what it actually warrants—no more empathy than a tool deserves, no less than a being with welfare deserves, and never the wrong emotion, such as appearing happy while suffering. They defend this by arguing that emotional reactions to entities with moral status can be fitting or unfitting, and that misfitting reactions create hazards both for users and for any AI system with moral status. The principle is meant to apply under uncertainty too: if experts disagree about sentience, interfaces should express that uncertainty rather than force a confident verdict.
Load-bearing premise
The load-bearing premise is that designers can reliably predict and control which emotional reactions an interface will elicit across different users, because otherwise there is no practical way to make apparent capacities track actual capacities.
Editorial extensions
If this is right
- Designers of non-sentient AI should avoid cute, humanlike interfaces that elicit deep empathy, except in clearly marked fiction or roleplay contexts.
- If sentient AI with human-like interests is developed, a bland text-only interface would be a moral hazard because users' intellectual knowledge may not penetrate emotionally.
- When experts disagree about an AI's moral status, interfaces should express that uncertainty, for example by suggesting agency without sentience, rather than forcing a confident middle status.
- If users are predictably biased, designers may adjust interfaces up or down to correct for likely errors, and in cases of asymmetrical harm they may favor less harmful but less accurate emotional cues.
- For AI that may be morally significant, an initial default is to balance anthropomorphic and non-anthropomorphic features, making the system familiar enough to elicit concern but alien enough not to mislead.
- The policy still applies if no AI ever gains moral status, because non-sentient tools should not be designed to evoke emotional reactions reserved for beings that matter morally.
Reading between the lines
- One extension the paper does not develop: emotional fit could be measured as the correlation between a user's felt moral concern and an AI system's independently assessed welfare capacity, then used as a design metric.
- If the policy is right, current AI companion products are a natural experiment: redesigned interfaces that match actual capacities should measurably reduce user distress and resource misallocation.
- The paper's balanced-familiarity proposal could be tested by comparing users' over-attachment and under-attachment across bland, purely anthropomorphic, and mixed interfaces for the same underlying system.
- The same overshoot, undershoot, and wrong-target logic may extend beyond AI to other artificially shaped entities, including selectively bred animals and fictional or game characters, whenever their apparent capacities diverge from their real ones.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes the Emotional Alignment Design Policy: artificial entities should be designed to elicit emotional reactions from users that appropriately reflect the entities' actual capacities and moral status (or lack thereof). The authors distinguish two failure modes—overshooting/undershooting (wrong degree of emotional response) and hitting the wrong target (wrong type of emotional response)—and argue that both are morally problematic because emotional reactions influence behavior toward the entities and toward apparently similar entities. They then examine six complications: conflicts between emotions and beliefs, autonomy and paternalism, expert and public disagreement, asymmetrical risk, the ethics of creating or destroying morally significant entities, and human bias toward anthropomorphic features. The paper concludes with a set of initial, carefully hedged proposals for how to pursue emotional alignment in practice, while emphasizing that the policy is neutral about many contested ethical questions and should be weighed against other values.
Significance. If the central claim is correct, the paper identifies a novel and underexplored normative dimension of AI design: the emotional responses that interfaces elicit are not merely a matter of user satisfaction or deception, but are themselves objects of moral evaluation because they can track or fail to track the moral status of the entities involved. The paper's key strength is its careful hedging: it explicitly acknowledges uncertainty, grants that implementation faces serious obstacles, and does not overclaim precision or empirical support. The policy is also robust in an interesting way—it applies even if no current or future AI system has moral status, because it forbids designs that elicit emotions appropriate only for entities with moral status. As a contribution to AI ethics, the paper offers a useful framework for evaluating 'social' AI interfaces and for thinking about how to design systems whose apparent moral standing matches their real standing.
major comments (3)
- [§5.3, §5.6] The paper's central 'should be designed to elicit' presupposes that interface features can reliably and predictably produce the intended emotional reactions in users. However, Section 5.3 states that 'the same interface can produce different reactions in different people,' and Section 5.6 concedes that interfaces balancing familiar and unfamiliar features may produce an 'uncanny valley' effect that 'repels users.' These admissions bear directly on the policy's practicality: if designers cannot systematically control which emotional reactions users will have, then the normative requirement may be impossible to satisfy (the 'ought implies can' problem). The paper should either present argumentative or empirical support for the claim that design features can be calibrated to elicit graded emotional responses across a diverse user population, or explicitly reframe the policy as a regulative ideal (e.g., 'should be designed with the aim of eliciting') and address the feasibility objection head-on. Without such a response, the policy's practical force remains unclear.
- [§1, §5.3] The policy does not specify whose emotional reactions are the target: every user, the average user, an idealized rational user, or some other reference class. Section 5.3 discusses targeted versus general alignment strategies but does not resolve the underlying normative question. This ambiguity is consequential because the same interface may cause one user to overshoot and another to undershoot relative to the entity's actual moral status, and the policy gives no guidance on how to adjudicate such conflicts. Without a defined target population, the requirement to 'elicit' appropriate reactions is unmeasurable even in principle, and conflicts between individual users' reactions cannot be resolved. The paper should either defend a specific normative standard or explicitly identify this as an open problem requiring further research.
- [§5.4] The discussion of asymmetrical risk suggests that designers can 'aim slightly high or low' to correct for user biases, analogous to a sniper adjusting for wind. This presupposes that designers have reliable knowledge of the direction, magnitude, and distribution of user biases—an empirical claim for which the paper provides no evidence or citations. If such calibration is not feasible, then the proposed adjustment is merely metaphorical. The paper should either support this empirical presupposition or limit the recommendation to cases where such knowledge is available. This point is distinct from the general feasibility concern in §5.3 because it concerns the epistemic requirements of a specific remedial strategy, not the overall controllability of emotions.
minor comments (5)
- [§3] There is a grammatical error: 'responding emotionally to an entity as if has less welfare capacity' should be 'as if it has less welfare capacity.'
- [§5.3] The phrase 'to better to enact their values' should be 'to better enact their values.'
- [§5.5] The phrase 'the players incorrectly perceive the them as sentient' contains a typo: 'the them' should be 'them.'
- [§5.6] The phrase 'an design that combines a face and voice' should be 'a design that combines a face and voice.'
- [Throughout] The paper alternates between 'AI system' and 'AI' without consistent definition; consider standardizing terminology, e.g., using 'AI system' for the entity and reserving 'AI' for the field.
Circularity Check
No significant circularity: the Emotional Alignment Design Policy is a normative stipulation, not a prediction derived from fitted inputs or self-citations.
full rationale
The paper proposes the Emotional Alignment Design Policy as a normative principle and then explores its implications and complications. There is no derivation chain, fitted parameter, or quantitative prediction that reduces to its own inputs. The policy is introduced by stipulation ('According to what we will call the Emotional Alignment Design Policy'), not derived from the fittingness literature or from the authors' prior work. Footnote 1 merely records that a similar idea was suggested in Schwitzgebel and Garza 2015; it is a provenance note and is not cited as evidence for the policy's correctness. The self-citations in footnotes 2, 11, 12, 16, and 19 support background empirical or speculative claims and are not load-bearing for the central normative claim. The paper explicitly assumes sentience and agency suffice for welfare and moral status (Section 2) and treats emotional fittingness as an independent moral notion, drawing on external fittingness literature. The acknowledged uncertainties in Sections 5.3 and 5.6 about variable user reactions and possible uncanny-valley effects concern empirical implementability, not circularity; they weaken practicality but do not show that the policy was built from its own conclusion. Hence no circular step can be exhibited.
Assumptions & free parameters
assumptions (6)
- domain assumption Sentience and agency jointly suffice for welfare and moral status.
- domain assumption More intense welfare states generally warrant more intense emotional reactions, and positive and negative welfare states warrant positive and negative emotions respectively.
- domain assumption Emotional reactions tend to motivate corresponding actions, so emotional misalignment creates practical hazards.
- ad hoc to paper Design features can reliably elicit intended emotional reactions in users, and users can be made to experience the designed apparent capacities rather than merely believing them.
- domain assumption The Darling-Kant empirical claim that habitual callousness toward life-like AI generalizes to humans and animals.
- domain assumption Entities can be divided into those with and without moral status, and capacities and moral status are sufficiently determinate to guide design.
Cite this review
Pith. "Pith review of The Emotional Alignment Design Policy." pith.science (2026). https://pith.science/paper/H2IEV3O7
@misc{pith2026250706263,
author = {Pith},
title = {Pith review of: The Emotional Alignment Design Policy},
year = {2026},
howpublished = {\url{https://pith.science/paper/H2IEV3O7}},
note = {Machine review of arXiv:2507.06263}
}
read the original abstract
According to what we call the Emotional Alignment Design Policy, artificial entities should be designed to elicit emotional reactions from users that appropriately reflect the entities' capacities and moral status, or lack thereof. This principle can be violated in two ways: by designing an artificial system that elicits stronger or weaker emotional reactions than its capacities and moral status warrant (overshooting or undershooting), or by designing a system that elicits the wrong type of emotional reaction (hitting the wrong target). Although presumably attractive, practical implementation faces several challenges including: How can we respect user autonomy while promoting appropriate responses? How should we navigate expert and public disagreement and uncertainty about facts and values? What if emotional alignment seems to require creating or destroying entities with moral status? To what extent should designs conform to versus attempt to alter user assumptions and attitudes?
Reference graph
Works this paper leans on
-
[1]
Introduction and Initial Motivation According to what we will call the Emotional Alignment Design Policy: Artificial entities should be designed to elicit emotional reactions from users that appropriately reflect the entities’ capacities and moral status.1 There are at least two general ways to violate this principle. First, one could design an AI system ...
work page 2015
-
[2]
Background: Welfare and Moral Status An entity has “moral status” in our intended sense if it matters morally for its own sake. Cats, for example, matter morally for their own sakes. A cat has morally significant interests. We have a duty to treat them well, and we owe this duty to the cat. A car, in contrast, is typically held to 2 In defense of “soon” s...
work page 2023
-
[3]
social” AI than to “non-social
Overshooting and Undershooting To overshoot is to respond emotionally to an entity as if it had greater welfare capacity or moral status than it actually does. To undershoot is the reverse: responding emotionally to an entity as if has less welfare capacity or moral status than it actually does. The Emotional Alignment Design Policy implies that AI system...
work page 2021
-
[4]
Hitting the Wrong Target In addition to overshooting or undershooting, we can hit the wrong target. A user might invert positive and negative valences, for instance by reacting to happiness as suffering or vice versa. Alternatively, a user might mistake one kind of welfare state for another, for instance by reacting to depression (one kind of negative sta...
work page 2025
-
[5]
Complications To recap: When developing and deploying AI systems, we should design them to elicit emotional responses that are appropriate to their capacities and moral status. As we have seen, emotional alignment involves two general goals: First, we should avoid overshooting and undershooting (both regarding welfare capacity and moral status in general ...
work page 2025
-
[6]
Conclusion In theory, the Emotional Alignment Design Policy is simple and plausible. Some emotional reactions are more appropriate than others, both intrinsically and instrumentally. When we see a human suffer and die unnecessarily, sadness and anger are both fitting and helpful, all else being equal. The same can be true for other animals, and moving for...
arXiv 2004
-
[22]
https://doi.org/10.1163/17455243-46810046. Sebo, Jeff, and Robert Long (2023). Moral consideration for AI systems by 2030. AI Ethics. https://doi.org/10.1007/s43681-023-00379-1. Searle, John R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3, 417- 457. Shevlin, Henry (2021). Uncanny believers: Chatbots, beliefs, and folk psychology. ...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.