REVIEW 3 major objections 5 minor 2 references
Ethics of generative AI and manipulation: a design-oriented research agenda
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This paper argues that good research on manipulation and generative AI depends on how 'manipulation' is defined, and defends the indifference criterion: manipulation is influence aimed at effectiveness but not explained by an aim to…
desk verdict Useful agenda-setting review of manipulation concepts for generative AI; the indifference criterion's non-intentional reading is underspecified exactly where the paper needs it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the indifference criterion for manipulation, proposed as a unified replacement for hidden-influence, bypassing-rationality, and trickery accounts. It has two parts: the influence must be aimed at a goal, excluding accidental influence, and the choice of the means of influence must not be explained by an aim to reveal reasons to the interlocutor. Because the criterion can be read functionally—asking what the means of influence is for—it applies to recommender systems and future persuasive large language models without relying on detectable intentions. It does the argument's work by turning manipulation from a question about hiddenness or psychological bypass into a question about the purpose embedded in how the influence was produced.
What would settle it
A controlled study could settle the sufficiency claim: present participants with an interface whose sign-up flow was A/B-tested purely for conversion, with no intention to reveal reasons, but with all persuasive elements fully disclosed; the indifference criterion predicts this counts as manipulation. If participants robustly judge it legitimate, or if a parallel study finds a clear case of manipulation where the influencer's chosen means was fully explained by the aim to reveal reasons, then the criterion would need revision.
Extended reading notes
Core claim
The central claim is that responsible design and regulation of generative AI should be organised around non-manipulation as a target value, and that the success of such design depends on the conceptualisation of manipulation chosen at the outset. Such a conceptualisation should satisfy a narrow criterion of capturing all and only cases of manipulation; the hidden-influence and bypassing-rationality accounts fail on both over- and under-inclusiveness, and disjunctive accounts multiply the same problems while obscuring what cases share. The paper's proposed criterion defines manipulation as an influence aimed at some goal whose chosen means is not explained by the aim to reveal reasons to the target. This makes it possible to describe the behaviour of AI systems as manipulative even when no human intended to trick anyone, since goals can be understood functionally rather than intentionally. The author acknowledges that the criterion needs further specification and operationalisation, and leaves open whether narrow or broad considerations should ultimately decide between conceptualisations.
Load-bearing premise
The argument rests on assuming that a single shared definition of manipulation exists and is the right tool for design, so that showing hidden-influence, bypassing-rationality, disjunctive, and trickery accounts fail is enough to license the indifference criterion.
Editorial extensions
If this is right
- Design for non-manipulation should focus on whether a system's output selection is explained by effectiveness rather than by revealing reasons, not on whether the influence is hidden or overt.
- Counterfactual checking becomes a candidate test: compare actual system output with the output a reason-revealing version of the system would give; divergence is evidence of indifference and hence manipulation.
- Fine-tuning large language models to optimise persuasive impact would count as manipulative design even if no individual intends to deceive, so alignment research needs to treat persuasive optimisation itself as ethically loaded.
- Human-feedback alignment methods are not a safe shortcut unless labelers' judgments can be shown to track the chosen conceptualisation; current evidence suggests folk judgments diverge from philosophical criteria.
- The criterion separates manipulation from deception: non-deceptive influence can still be manipulative, so regulation aimed only at false content misses a core risk.
Reading between the lines
- As an extension the paper leaves implicit, one testable proxy for indifference is to prompt the same model to explain its reasoning to the user and compare that baseline output with production outputs; divergence could indicate manipulation at scale.
- If folk concepts are as context-dependent as the experiments the paper cites suggest, the monistic indifference criterion may need to be supplemented by situation-specific operationalisations even if a single definition remains the theoretical target.
- The framework offers a principled way to redraw the nudge boundary: a nudge is non-manipulative when its function is to help the target see reasons, and manipulative when its function is only effectiveness; this is a normative test that behavioural policy could adopt.
- The paper's design focus leaves the regulatory question open; a natural next question is whether the indifference criterion can be turned into auditable system specifications, since intention-based criteria cannot be directly inspected in deployed models.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper defends a design-oriented research agenda for non-manipulative generative AI. Its central thesis is that research on manipulation and generative AI depends significantly on the conceptualisation of manipulation adopted. The author reviews and rejects four alternative criteria—hidden influence, bypassing rationality, disjunctive conceptions, and trickery—and endorses the indifference criterion, according to which manipulation is influence that aims to be effective but is not explained by the aim to reveal reasons to the interlocutor. The paper embeds this endorsement in a design-for-values framework, proposing conceptual, empirical, and design-stage research questions, and it explicitly acknowledges unresolved issues in specifying and operationalising the indifference criterion.
Significance. If the argument succeeds, the paper makes a useful contribution by directing attention to the conceptual foundations of AI-manipulation research and by connecting philosophical analysis to design requirements, a connection that is often missing in AI ethics discussions such as Weidinger et al. (2022). Its strengths include a clear and honest statement of the open problems of the indifference criterion (specification of the ideal state and operationalisation), explicit engagement with the hidden-influence and bypassing-rationality literatures, and a structured presentation of empirical and design questions. The paper is also transparent about the limitations of a research agenda, repeatedly flagging what it does not settle. However, the load-bearing application of the criterion to current LLMs remains indeterminate, which currently weakens the claim that the indifference criterion is the most appropriate conceptualisation for emergent manipulation.
major comments (3)
- [The indifference criterion] The application of the indifference criterion to current LLMs is indeterminate in a load-bearing way. The paper states that ChatGPT-like systems are 'not yet capable of fine-tuning their output in pursuit of goals other than text-sequence prediction' (p. 9, n. 24), yet it claims the criterion's chief advantage is capturing emergent manipulation exemplified by stochastic parrots. If 'aims to be effective' is read intentionalistically, a next-token predictor does not aim at effective influence, so the stochastic-parrots case is not captured. If it is read functionally as whatever the system is optimised for, then text-sequence prediction is the relevant function and ordinary LLM outputs are not manipulative either; if, instead, any engagement effect counts as an aim, the criterion over-generates and the acknowledged risk that 'generative AI systems come away as necessarily manipulative' (p. 9) becomes actual. The manuscript needs a principled account of how 'aim' is attributed to AI systems in functional terms, one that distinguishes next-token prediction from effective influence, before the claimed advantage over the trickery criterion is established.
- [The indifference criterion] The claim that the indifference criterion 'fares well on the narrow criterion of appropriateness' is not yet supported. The narrow criterion requires capturing all and only cases of manipulation, and the paper rejects alternative criteria for over- and under-inclusiveness. However, the indifference criterion's own key constituents—'aims to be effective,' 'revealing reasons to the interlocutor,' and the 'ideal state'—remain unspecified; the paper explicitly concedes that the ideal state needs to be specified in more detail and that the criterion 'must be further specified and operationalised.' Without such specification, the extension of the criterion is indeterminate, so its asserted superiority over hidden-influence and bypassing-rationality criteria cannot be assessed. Because the central claim is that conceptualisation choice is foundational for design, this gap is load-bearing.
- [Disjunctive conceptions of manipulation] The argument against disjunctive criteria assumes that the absence of a common factor behind all forms of manipulation is a decisive theoretical cost. This assumption is precisely what the pluralist literature, including Noggle (2022) and the possibility noted by the paper that 'there are simply different types of manipulation' (p. 7), denies. The paper does not explain why design for non-manipulation could not proceed by specifying separate requirements for each disjunct (a consequence it calls a practical problem) nor why the monistic indifference criterion is preferable to a pluralist research program. Since the paper's central claim is that one appropriate conceptualisation can guide non-manipulative design, the monistic assumption needs an explicit defence or qualification.
minor comments (5)
- [Introduction] The sentence 'The section “Design for values and conceptual engineering” “Design for non-manipulation”' appears to have a formatting or drafting error; the section reference is garbled.
- [Design stage] On page 11, 'human souls classify sample outputs' should read 'human labelers classify sample outputs'.
- [The trickery criterion] On page 8, 'it generative AI threatens' is missing a word; it should likely read 'when generative AI threatens'.
- [The indifference criterion] Footnote 22 is inconsistent in its citation of the author's own work: it mentions 'Klenk (2021)' where the reference list and surrounding text use 'Klenk (2021c)'.
- [Empirical stage] The claim that Osman and Bechlivanidis are 'the only ones that explicitly address folk-conceptions of manipulation' (p. 10) is strong and would benefit from a broader citation base, since it is hard to verify and may be read as overclaiming.
Circularity Check
No significant circularity: the paper's positive case for the indifference criterion is argued in-text against competing accounts, with self-citations serving as pointers to prior elaboration rather than as the sole load-bearing justification.
full rationale
This is a philosophy research-agenda paper, not a formal derivation, so the circularity tests must be applied to whether the conclusion is presupposed by its own inputs. The central claim is that the indifference criterion is the most appropriate conceptualisation of manipulation for designing non-manipulative generative AI. The criterion is indeed drawn from the author's prior work (Klenk 2020, 2021c, 2022a, 2022b, 2023), and several key moves cite those works. However, the paper does not simply assert the criterion; it argues for it by rejecting the hidden-influence, bypassing-rationality, disjunctive, and trickery accounts on stated grounds (over- and under-inclusiveness, practical and ethical costs, failure to handle emergent non-intentional manipulation), and it gives an independent rationale for the indifference criterion's advantage: it can be interpreted non-intentionally via the function of a chosen means of influence, thereby capturing emergent manipulation by generative AI. The paper also acknowledges unresolved issues, including the need to specify the ideal state and to operationalise the criterion, and it flags the risk that generative AI systems come away as necessarily manipulative. These are substantive philosophical limitations, not circular reductions. There are self-citations, but they function as pointers to more detailed prior treatments; the in-text comparative arguments carry the weight. No equation is fitted and then renamed as a prediction, no uniqueness theorem is imported from the authors, and no ansatz is smuggled in solely through a self-citation. The closest candidate for a circularity concern is that the criterion is defined so as to include non-intentional functional influence and then praised for capturing non-intentional emergent manipulation, but that is a normal feature of conceptual analysis, not a circular derivation. Overall, the derivation chain is self-contained in the relevant sense: the conclusion is substantively argued against alternatives, and the acknowledged open questions concern operationalisation, not a disguised restatement of the inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption There exists an appropriate conceptualisation of manipulation that captures all and only cases of manipulation, the narrow criterion of appropriateness.
- domain assumption Manipulation is best understood as a form of influence aimed at a goal, not as accidental influence.
- domain assumption Design for values can translate a target value, such as non-manipulation, into concrete design requirements.
Cite this review
Pith. "Pith review of Ethics of generative AI and manipulation: a design-oriented research agenda." pith.science (2026). https://pith.science/paper/YZ6FDF5E
@misc{pith2026250304733,
author = {Pith},
title = {Pith review of: Ethics of generative AI and manipulation: a design-oriented research agenda},
year = {2026},
howpublished = {\url{https://pith.science/paper/YZ6FDF5E}},
note = {Machine review of arXiv:2503.04733}
}
read the original abstract
Generative AI enables automated, effective manipulation at scale. Despite the growing general ethical discussion around generative AI, the specific manipulation risks remain inadequately investigated. This article outlines essential inquiries encompassing conceptual, empirical, and design dimensions of manipulation, pivotal for comprehending and curbing manipulation risks. By highlighting these questions, the article underscores the necessity of an appropriate conceptualisation of manipulation to ensure the responsible development of Generative AI technologies.
Reference graph
Works this paper leans on
-
[10]
2307/ 32197 40 Baron, M. (2014). The mens rea and moral status of manipulation. In C. Coons & M. Weber (Eds.), Manipulation: Theory and practice (pp. 98–109). Oxford University Press. Beauchamp, T. L. (1984). Manipulative advertising. Business and Pro- fessional Ethics Journal, 3, 1–22. Beauchamp, T. L., & Childress, J. F. (2019). Principles of biomedical...
work page 2014
-
[4569]
https:// doi. org/ 10. 1038/ s41598- 023- 31341-0 Matz, S., Teeny, J., Vaid, S. S., Harari, G. M., & Cerf, M. (2023). The potential of generative AI for personalized persuasion at scale. https:// doi. org/ 10. 31234/ osf. io/ rn97c Mills, C. (1995). Politics and manipulation. Social Theory and Prac- tice, 21(1), 97–112. Noggle, R. (1996). Manipulative act...
work page 2023
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.