Pith. sign in

REVIEW 3 major objections 5 minor 2 references

Ethics of generative AI and manipulation: a design-oriented research agenda

T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This paper argues that good research on manipulation and generative AI depends on how 'manipulation' is defined, and defends the indifference criterion: manipulation is influence aimed at effectiveness but not explained by an aim to…

desk verdict Useful agenda-setting review of manipulation concepts for generative AI; the indifference criterion's non-intentional reading is underspecified exactly where the paper needs it. read the letter →

arxiv 2503.04733 v1 pith:YZ6FDF5E submitted 2025-02-01 cs.CY cs.AI

classification cs.CYcs.AI
keywords generativeAImanipulationlargelanguagemodelsvaluesensitivedesignethicsindifferencecriterionconceptualengineeringpersuasion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the ethical debate about manipulation in generative AI cannot move forward until researchers pick an appropriate definition of 'manipulation', because every definition acts like a searchlight: it reveals some phenomena and hides others, and different definitions lead to different design requirements and different regulations. It reviews the leading candidates—hidden influence, bypassing rationality, disjunctive criteria, and trickery—and finds each either over- or under-inclusive for the cases generative AI makes salient. It then defends the indifference criterion: manipulation is influence aimed at effectiveness that is not explained by an aim to reveal reasons to the person influenced. On this definition, intentional fraud, unwitting manipulation through A/B-tested dark patterns, and emergent non-intentional manipulation by AI systems all count as manipulation, while many hidden or emotion-appealing influences do not. The paper is a design-oriented research agenda, so it frames the work still needed—empirical study of folk concepts, operationalisation of the criterion, and translation into concrete design requirements—rather than a finished solution.

What carries the argument

The central machinery is the indifference criterion for manipulation, proposed as a unified replacement for hidden-influence, bypassing-rationality, and trickery accounts. It has two parts: the influence must be aimed at a goal, excluding accidental influence, and the choice of the means of influence must not be explained by an aim to reveal reasons to the interlocutor. Because the criterion can be read functionally—asking what the means of influence is for—it applies to recommender systems and future persuasive large language models without relying on detectable intentions. It does the argument's work by turning manipulation from a question about hiddenness or psychological bypass into a question about the purpose embedded in how the influence was produced.

What would settle it

A controlled study could settle the sufficiency claim: present participants with an interface whose sign-up flow was A/B-tested purely for conversion, with no intention to reveal reasons, but with all persuasive elements fully disclosed; the indifference criterion predicts this counts as manipulation. If participants robustly judge it legitimate, or if a parallel study finds a clear case of manipulation where the influencer's chosen means was fully explained by the aim to reveal reasons, then the criterion would need revision.

Watch

Extended reading notes

Core claim

The central claim is that responsible design and regulation of generative AI should be organised around non-manipulation as a target value, and that the success of such design depends on the conceptualisation of manipulation chosen at the outset. Such a conceptualisation should satisfy a narrow criterion of capturing all and only cases of manipulation; the hidden-influence and bypassing-rationality accounts fail on both over- and under-inclusiveness, and disjunctive accounts multiply the same problems while obscuring what cases share. The paper's proposed criterion defines manipulation as an influence aimed at some goal whose chosen means is not explained by the aim to reveal reasons to the target. This makes it possible to describe the behaviour of AI systems as manipulative even when no human intended to trick anyone, since goals can be understood functionally rather than intentionally. The author acknowledges that the criterion needs further specification and operationalisation, and leaves open whether narrow or broad considerations should ultimately decide between conceptualisations.

Load-bearing premise

The argument rests on assuming that a single shared definition of manipulation exists and is the right tool for design, so that showing hidden-influence, bypassing-rationality, disjunctive, and trickery accounts fail is enough to license the indifference criterion.

Editorial extensions

If this is right

  • Design for non-manipulation should focus on whether a system's output selection is explained by effectiveness rather than by revealing reasons, not on whether the influence is hidden or overt.
  • Counterfactual checking becomes a candidate test: compare actual system output with the output a reason-revealing version of the system would give; divergence is evidence of indifference and hence manipulation.
  • Fine-tuning large language models to optimise persuasive impact would count as manipulative design even if no individual intends to deceive, so alignment research needs to treat persuasive optimisation itself as ethically loaded.
  • Human-feedback alignment methods are not a safe shortcut unless labelers' judgments can be shown to track the chosen conceptualisation; current evidence suggests folk judgments diverge from philosophical criteria.
  • The criterion separates manipulation from deception: non-deceptive influence can still be manipulative, so regulation aimed only at false content misses a core risk.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • As an extension the paper leaves implicit, one testable proxy for indifference is to prompt the same model to explain its reasoning to the user and compare that baseline output with production outputs; divergence could indicate manipulation at scale.
  • If folk concepts are as context-dependent as the experiments the paper cites suggest, the monistic indifference criterion may need to be supplemented by situation-specific operationalisations even if a single definition remains the theoretical target.
  • The framework offers a principled way to redraw the nudge boundary: a nudge is non-manipulative when its function is to help the target see reasons, and manipulative when its function is only effectiveness; this is a normative test that behavioural policy could adopt.
  • The paper's design focus leaves the regulatory question open; a natural next question is whether the indifference criterion can be turned into auditable system specifications, since intention-based criteria cannot be directly inspected in deployed models.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper defends a design-oriented research agenda for non-manipulative generative AI. Its central thesis is that research on manipulation and generative AI depends significantly on the conceptualisation of manipulation adopted. The author reviews and rejects four alternative criteria—hidden influence, bypassing rationality, disjunctive conceptions, and trickery—and endorses the indifference criterion, according to which manipulation is influence that aims to be effective but is not explained by the aim to reveal reasons to the interlocutor. The paper embeds this endorsement in a design-for-values framework, proposing conceptual, empirical, and design-stage research questions, and it explicitly acknowledges unresolved issues in specifying and operationalising the indifference criterion.

Significance. If the argument succeeds, the paper makes a useful contribution by directing attention to the conceptual foundations of AI-manipulation research and by connecting philosophical analysis to design requirements, a connection that is often missing in AI ethics discussions such as Weidinger et al. (2022). Its strengths include a clear and honest statement of the open problems of the indifference criterion (specification of the ideal state and operationalisation), explicit engagement with the hidden-influence and bypassing-rationality literatures, and a structured presentation of empirical and design questions. The paper is also transparent about the limitations of a research agenda, repeatedly flagging what it does not settle. However, the load-bearing application of the criterion to current LLMs remains indeterminate, which currently weakens the claim that the indifference criterion is the most appropriate conceptualisation for emergent manipulation.

major comments (3)
  1. [The indifference criterion] The application of the indifference criterion to current LLMs is indeterminate in a load-bearing way. The paper states that ChatGPT-like systems are 'not yet capable of fine-tuning their output in pursuit of goals other than text-sequence prediction' (p. 9, n. 24), yet it claims the criterion's chief advantage is capturing emergent manipulation exemplified by stochastic parrots. If 'aims to be effective' is read intentionalistically, a next-token predictor does not aim at effective influence, so the stochastic-parrots case is not captured. If it is read functionally as whatever the system is optimised for, then text-sequence prediction is the relevant function and ordinary LLM outputs are not manipulative either; if, instead, any engagement effect counts as an aim, the criterion over-generates and the acknowledged risk that 'generative AI systems come away as necessarily manipulative' (p. 9) becomes actual. The manuscript needs a principled account of how 'aim' is attributed to AI systems in functional terms, one that distinguishes next-token prediction from effective influence, before the claimed advantage over the trickery criterion is established.
  2. [The indifference criterion] The claim that the indifference criterion 'fares well on the narrow criterion of appropriateness' is not yet supported. The narrow criterion requires capturing all and only cases of manipulation, and the paper rejects alternative criteria for over- and under-inclusiveness. However, the indifference criterion's own key constituents—'aims to be effective,' 'revealing reasons to the interlocutor,' and the 'ideal state'—remain unspecified; the paper explicitly concedes that the ideal state needs to be specified in more detail and that the criterion 'must be further specified and operationalised.' Without such specification, the extension of the criterion is indeterminate, so its asserted superiority over hidden-influence and bypassing-rationality criteria cannot be assessed. Because the central claim is that conceptualisation choice is foundational for design, this gap is load-bearing.
  3. [Disjunctive conceptions of manipulation] The argument against disjunctive criteria assumes that the absence of a common factor behind all forms of manipulation is a decisive theoretical cost. This assumption is precisely what the pluralist literature, including Noggle (2022) and the possibility noted by the paper that 'there are simply different types of manipulation' (p. 7), denies. The paper does not explain why design for non-manipulation could not proceed by specifying separate requirements for each disjunct (a consequence it calls a practical problem) nor why the monistic indifference criterion is preferable to a pluralist research program. Since the paper's central claim is that one appropriate conceptualisation can guide non-manipulative design, the monistic assumption needs an explicit defence or qualification.
minor comments (5)
  1. [Introduction] The sentence 'The section “Design for values and conceptual engineering” “Design for non-manipulation”' appears to have a formatting or drafting error; the section reference is garbled.
  2. [Design stage] On page 11, 'human souls classify sample outputs' should read 'human labelers classify sample outputs'.
  3. [The trickery criterion] On page 8, 'it generative AI threatens' is missing a word; it should likely read 'when generative AI threatens'.
  4. [The indifference criterion] Footnote 22 is inconsistent in its citation of the author's own work: it mentions 'Klenk (2021)' where the reference list and surrounding text use 'Klenk (2021c)'.
  5. [Empirical stage] The claim that Osman and Bechlivanidis are 'the only ones that explicitly address folk-conceptions of manipulation' (p. 10) is strong and would benefit from a broader citation base, since it is hard to verify and may be read as overclaiming.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the paper's positive case for the indifference criterion is argued in-text against competing accounts, with self-citations serving as pointers to prior elaboration rather than as the sole load-bearing justification.

full rationale

This is a philosophy research-agenda paper, not a formal derivation, so the circularity tests must be applied to whether the conclusion is presupposed by its own inputs. The central claim is that the indifference criterion is the most appropriate conceptualisation of manipulation for designing non-manipulative generative AI. The criterion is indeed drawn from the author's prior work (Klenk 2020, 2021c, 2022a, 2022b, 2023), and several key moves cite those works. However, the paper does not simply assert the criterion; it argues for it by rejecting the hidden-influence, bypassing-rationality, disjunctive, and trickery accounts on stated grounds (over- and under-inclusiveness, practical and ethical costs, failure to handle emergent non-intentional manipulation), and it gives an independent rationale for the indifference criterion's advantage: it can be interpreted non-intentionally via the function of a chosen means of influence, thereby capturing emergent manipulation by generative AI. The paper also acknowledges unresolved issues, including the need to specify the ideal state and to operationalise the criterion, and it flags the risk that generative AI systems come away as necessarily manipulative. These are substantive philosophical limitations, not circular reductions. There are self-citations, but they function as pointers to more detailed prior treatments; the in-text comparative arguments carry the weight. No equation is fitted and then renamed as a prediction, no uniqueness theorem is imported from the authors, and no ansatz is smuggled in solely through a self-citation. The closest candidate for a circularity concern is that the criterion is defined so as to include non-intentional functional influence and then praised for capturing non-intentional emergent manipulation, but that is a normal feature of conceptual analysis, not a circular derivation. Overall, the derivation chain is self-contained in the relevant sense: the conclusion is substantively argued against alternatives, and the acknowledged open questions concern operationalisation, not a disguised restatement of the inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper does not introduce new empirical free parameters or invented entities. Its structural assumptions are the monistic-conceptualisation premise, the goal-directedness premise, and the design-for-values framework. The author explicitly flags open questions about the first and third of these premises.

assumptions (3)
  • domain assumption There exists an appropriate conceptualisation of manipulation that captures all and only cases of manipulation, the narrow criterion of appropriateness.
    The paper treats the narrow criterion as a legitimate benchmark and uses it to reject hidden influence, bypassing rationality, and disjunctive accounts. If no monistic concept satisfies this criterion, the argument's framework would need revision.
  • domain assumption Manipulation is best understood as a form of influence aimed at a goal, not as accidental influence.
    The paper adopts this restriction following the cited manipulation literature, and it excludes accidental influence without arguing for the restriction.
  • domain assumption Design for values can translate a target value, such as non-manipulation, into concrete design requirements.
    The paper relies on the design-for-values literature (van de Poel, van den Hoven, Veluwenkamp) as a viable methodology for translating the conceptualisation into design practice.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Ethics of generative AI and manipulation: a design-oriented research agenda." pith.science (2026). https://pith.science/paper/YZ6FDF5E

@misc{pith2026250304733,
  author       = {Pith},
  title        = {Pith review of: Ethics of generative AI and manipulation: a design-oriented research agenda},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YZ6FDF5E}},
  note         = {Machine review of arXiv:2503.04733}
}
read the original abstract

Generative AI enables automated, effective manipulation at scale. Despite the growing general ethical discussion around generative AI, the specific manipulation risks remain inadequately investigated. This article outlines essential inquiries encompassing conceptual, empirical, and design dimensions of manipulation, pivotal for comprehending and curbing manipulation risks. By highlighting these questions, the article underscores the necessity of an appropriate conceptualisation of manipulation to ensure the responsible development of Generative AI technologies.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

2 extracted references · 2 canonical work pages

  1. [10]

    2307/ 32197 40 Baron, M. (2014). The mens rea and moral status of manipulation. In C. Coons & M. Weber (Eds.), Manipulation: Theory and practice (pp. 98–109). Oxford University Press. Beauchamp, T. L. (1984). Manipulative advertising. Business and Pro- fessional Ethics Journal, 3, 1–22. Beauchamp, T. L., & Childress, J. F. (2019). Principles of biomedical...

  2. [4569]

    https:// doi. org/ 10. 1038/ s41598- 023- 31341-0 Matz, S., Teeny, J., Vaid, S. S., Harari, G. M., & Cerf, M. (2023). The potential of generative AI for personalized persuasion at scale. https:// doi. org/ 10. 31234/ osf. io/ rn97c Mills, C. (1995). Politics and manipulation. Social Theory and Prac- tice, 21(1), 97–112. Noggle, R. (1996). Manipulative act...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.