{"id":"5d163dd7-3f8d-4ccd-99b6-055e125a77e2","arxiv_id":"2503.04733","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The article presents a research agenda arguing that an adequate conceptualization of manipulation is a prerequisite for designing non-manipulative generative AI.","lead":"This paper argues that designing generative AI so it does not manipulate people depends on first agreeing what manipulation means, and compares several philosophical definitions. It maps out the conceptual, empirical, and design questions a research program on non-manipulative AI should tackle.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The indifference criterion's non-intentional reading is indeterminate for current LLMs: if outputs 'aim' only at next-token prediction, the stochastic-parrot case is not captured; if any engagement effect counts, the criterion over-generates and design guidance collapses.","rationale":"The reader's weakest assumption—that the argument presupposes a monistic conceptualisation of manipulation—is a genuine high-level concern, and I partially agree. But my stress-test identifies a more concrete, internal gap: the indifference criterion's non-intentional interpretation is underdetermined when applied to current LLMs. This directly threatens the claimed advantage over the trickery criterion, since emergent manipulation is the central case the indifference view is supposed to capture. The paper openly acknowledges the need for further specification and operationalisation, and it even notes the risk that generative AI comes away as necessarily manipulative. Given that the article is explicitly a research agenda rather than a full defense of the criterion, these open questions are consistent with a conditional acceptance. The concern does not warrant rejection because the paper's primary thesis—that conceptualisation matters for design—survives even if the indifference criterion needs refinement; the agenda remains valuable. I therefore recommend no change to the reader's CONDITIONAL verdict, but I would emphasize that the functional-aim question, not just monism, is a key condition that must be settled before the indifference criterion can ground design requirements.","tokens_in":20261,"tokens_out":5199,"duration_ms":57981,"concrete_test":"Formalize the indifference criterion for a generative AI system as follows: let O be the system's objective function, and define 'aims to be effective' as 'outputs are selected to optimize O.' Consider two systems: (a) GPT-3 fine-tuned only with next-token log-loss, and (b) the same architecture fine-tuned to maximize user engagement. Apply the criterion to an identical user-influencing output from each system. If (a) is non-manipulative and (b) is manipulative, then the criterion depends on the training objective and one must justify why next-token prediction is not an 'aim to be effective' while engagement is. If both are manipulative, then the criterion classifies nearly all deployed LLMs as manipulative, and the design-for-non-manipulation goal becomes vacuous. If both are non-manipulative, the stochastic-parrot case in section 'The indifference criterion' is not captured.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the indifference criterion is the most appropriate conceptualisation because, unlike trickery, it captures emergent, non-intentional manipulation by generative AI. The criterion is stated as: manipulation is an influence that aims to be effective but is not explained by the aim to reveal reasons to the interlocutor (section 'The indifference criterion'). The paper explicitly allows a non-intentional interpretation by appealing to the function of a chosen means of influence, e.g., a recommender system's 'watch next video' choice has the function to induce target behavior. This is the load-bearing move: without a determinate account of functional aims, the criterion cannot reliably classify current generative AI systems. Take an LLM trained only for next-token prediction (e.g., GPT-3 before RLHF). Its outputs may influence users, but its objective function is text-sequence prediction, not 'effective influence.' If we read the first condition strictly, such outputs do not aim to be effective, so they are not manipulative—contradicting the paper's stochastic-parrots/bullshit analogy and failing to capture the very emergent manipulation the criterion was designed to catch. If we instead read 'aim' functionally as whatever the system is optimized to achieve, then any engagement-optimized or persuasion-optimized system counts as manipulative, and the paper's acknowledged risk that 'generative AI systems come away as necessarily manipulative' becomes a real over-inclusiveness problem. The paper flags this risk but does not resolve it. Since the indifference criterion's advantage over trickery depends on handling exactly this non-intentional case, the underspecification is not a minor operational detail: it is a gap in the argument that the criterion is the most appropriate conceptualisation for design.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper defends a design-oriented research agenda for non-manipulative generative AI. Its central thesis is that research on manipulation and generative AI depends significantly on the conceptualisation of manipulation adopted. The author reviews and rejects four alternative criteria—hidden influence, bypassing rationality, disjunctive conceptions, and trickery—and endorses the indifference criterion, according to which manipulation is influence that aims to be effective but is not explained by the aim to reveal reasons to the interlocutor. The paper embeds this endorsement in a design-for-values framework, proposing conceptual, empirical, and design-stage research questions, and it explicitly acknowledges unresolved issues in specifying and operationalising the indifference criterion.","tokens_in":20472,"tokens_out":7542,"duration_ms":67295,"significance":"If the argument succeeds, the paper makes a useful contribution by directing attention to the conceptual foundations of AI-manipulation research and by connecting philosophical analysis to design requirements, a connection that is often missing in AI ethics discussions such as Weidinger et al. (2022). Its strengths include a clear and honest statement of the open problems of the indifference criterion (specification of the ideal state and operationalisation), explicit engagement with the hidden-influence and bypassing-rationality literatures, and a structured presentation of empirical and design questions. The paper is also transparent about the limitations of a research agenda, repeatedly flagging what it does not settle. However, the load-bearing application of the criterion to current LLMs remains indeterminate, which currently weakens the claim that the indifference criterion is the most appropriate conceptualisation for emergent manipulation.","major_comments":[{"comment":"The application of the indifference criterion to current LLMs is indeterminate in a load-bearing way. The paper states that ChatGPT-like systems are 'not yet capable of fine-tuning their output in pursuit of goals other than text-sequence prediction' (p. 9, n. 24), yet it claims the criterion's chief advantage is capturing emergent manipulation exemplified by stochastic parrots. If 'aims to be effective' is read intentionalistically, a next-token predictor does not aim at effective influence, so the stochastic-parrots case is not captured. If it is read functionally as whatever the system is optimised for, then text-sequence prediction is the relevant function and ordinary LLM outputs are not manipulative either; if, instead, any engagement effect counts as an aim, the criterion over-generates and the acknowledged risk that 'generative AI systems come away as necessarily manipulative' (p. 9) becomes actual. The manuscript needs a principled account of how 'aim' is attributed to AI systems in functional terms, one that distinguishes next-token prediction from effective influence, before the claimed advantage over the trickery criterion is established.","section":"The indifference criterion"},{"comment":"The claim that the indifference criterion 'fares well on the narrow criterion of appropriateness' is not yet supported. The narrow criterion requires capturing all and only cases of manipulation, and the paper rejects alternative criteria for over- and under-inclusiveness. However, the indifference criterion's own key constituents—'aims to be effective,' 'revealing reasons to the interlocutor,' and the 'ideal state'—remain unspecified; the paper explicitly concedes that the ideal state needs to be specified in more detail and that the criterion 'must be further specified and operationalised.' Without such specification, the extension of the criterion is indeterminate, so its asserted superiority over hidden-influence and bypassing-rationality criteria cannot be assessed. Because the central claim is that conceptualisation choice is foundational for design, this gap is load-bearing.","section":"The indifference criterion"},{"comment":"The argument against disjunctive criteria assumes that the absence of a common factor behind all forms of manipulation is a decisive theoretical cost. This assumption is precisely what the pluralist literature, including Noggle (2022) and the possibility noted by the paper that 'there are simply different types of manipulation' (p. 7), denies. The paper does not explain why design for non-manipulation could not proceed by specifying separate requirements for each disjunct (a consequence it calls a practical problem) nor why the monistic indifference criterion is preferable to a pluralist research program. Since the paper's central claim is that one appropriate conceptualisation can guide non-manipulative design, the monistic assumption needs an explicit defence or qualification.","section":"Disjunctive conceptions of manipulation"}],"minor_comments":[{"comment":"The sentence 'The section “Design for values and conceptual engineering” “Design for non-manipulation”' appears to have a formatting or drafting error; the section reference is garbled.","section":"Introduction"},{"comment":"On page 11, 'human souls classify sample outputs' should read 'human labelers classify sample outputs'.","section":"Design stage"},{"comment":"On page 8, 'it generative AI threatens' is missing a word; it should likely read 'when generative AI threatens'.","section":"The trickery criterion"},{"comment":"Footnote 22 is inconsistent in its citation of the author's own work: it mentions 'Klenk (2021)' where the reference list and surrounding text use 'Klenk (2021c)'.","section":"The indifference criterion"},{"comment":"The claim that Osman and Bechlivanidis are 'the only ones that explicitly address folk-conceptions of manipulation' (p. 10) is strong and would benefit from a broader citation base, since it is hard to verify and may be read as overclaiming.","section":"Empirical stage"}],"recommendation":"major_revision","confidential_remarks":"The paper's endorsement of the indifference criterion is continuous with the author's own prior publications, and the positive verdict is consistent with that earlier work. This is not circular in a formal sense, but given the paper's centrality to the evaluation, I recommend that the editor seek a reviewer with expertise in the philosophy of manipulation (e.g., on Noggle's pluralist objections) who is not connected to the author. The fit with Ethics and Information Technology is appropriate for a research-agenda paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe short version: this is a competent, honest agenda-setting paper, and the comparative review of manipulation criteria is genuinely useful. But it is not a new result, and the central positive claim—that the indifference criterion is the most appropriate conceptualisation—has a soft spot where it matters most.\n\nWhat the paper does well: it situates the conceptual question inside a design-for-values framework and shows why the choice of conceptualisation has practical consequences for design and regulation. The critiques of hidden influence, bypassing rationality, and disjunctive accounts are fair and well referenced. The research questions for the empirical and design stages are sensible, and the paper is unusually candid about its own open problems, especially the need to operationalise 'indifference' and specify the ideal state. Much of the positive account is drawn from the author's own prior work, so the novelty is in the agenda packaging rather than in the criterion itself—not a flaw, but worth knowing before citing it as an independent source.\n\nThe main weakness is the non-intentional reading of the indifference criterion. The criterion is stated as influence aimed at effectiveness but not explained by the aim to reveal reasons to the interlocutor. The paper says this can be read functionally, via the function of the chosen means of influence, and that this lets us capture emergent, unwitting manipulation by generative AI. But the functional reading is not worked out. Consider a current LLM trained only for next-token prediction. Its outputs influence users, but its objective is text-sequence prediction, not effective influence. On a strict reading, it fails the first condition and is not manipulative—which undercuts the paper's stochastic-parrots analogy. If 'aim' is read functionally as whatever the system is optimised to achieve, then any engagement-optimised system counts as manipulative, and the paper's own flag that generative AI might come out as 'necessarily manipulative' becomes a real over-inclusiveness problem, not a minor caveat. The paper acknowledges all this but does not resolve it. Since the indifference criterion's advantage over the trickery account depends precisely on handling this non-intentional case, the underspecification is a genuine gap in the argument.\n\nThere is also a monism assumption: the paper assumes a single, unified conceptualisation is both possible and preferable for design purposes. Noggle's work, which the paper cites, contains serious objections to that assumption, and the paper does not engage with them beyond a nod. Again, acknowledged but unresolved.\n\nThese are problems for someone who wants to adopt the indifference criterion as a working definition. As a research agenda, the paper is honest about what remains open; it is neither sloppy nor overclaimed. I would send it to review—it deserves serious engagement, and the gaps I name are productive ones for the field. If you write about manipulation and generative AI, cite it as a map, but don't treat the criterion as settled.\n\nBest,\n[You]","headline":"Useful agenda-setting review of manipulation concepts for generative AI; the indifference criterion's non-intentional reading is underspecified exactly where the paper needs it.","tokens_in":21088,"tokens_out":3804,"would_cite":true,"duration_ms":32276,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that good research on manipulation and generative AI depends on how 'manipulation' is defined, and defends the indifference criterion: manipulation is influence aimed at effectiveness but not explained by an aim to…","keywords":["generative AI","manipulation","large language models","value sensitive design","AI ethics","indifference criterion","conceptual engineering","persuasion"],"falsifier":"A controlled study could settle the sufficiency claim: present participants with an interface whose sign-up flow was A/B-tested purely for conversion, with no intention to reveal reasons, but with all persuasive elements fully disclosed; the indifference criterion predicts this counts as manipulation. If participants robustly judge it legitimate, or if a parallel study finds a clear case of manipulation where the influencer's chosen means was fully explained by the aim to reveal reasons, then the criterion would need revision.","tokens_in":20011,"feed_emoji":"🤖","tokens_out":6514,"duration_ms":61496,"temperature":0.7,"pith_summary":"This paper argues that the ethical debate about manipulation in generative AI cannot move forward until researchers pick an appropriate definition of 'manipulation', because every definition acts like a searchlight: it reveals some phenomena and hides others, and different definitions lead to different design requirements and different regulations. It reviews the leading candidates—hidden influence, bypassing rationality, disjunctive criteria, and trickery—and finds each either over- or under-inclusive for the cases generative AI makes salient. It then defends the indifference criterion: manipulation is influence aimed at effectiveness that is not explained by an aim to reveal reasons to the person influenced. On this definition, intentional fraud, unwitting manipulation through A/B-tested dark patterns, and emergent non-intentional manipulation by AI systems all count as manipulation, while many hidden or emotion-appealing influences do not. The paper is a design-oriented research agenda, so it frames the work still needed—empirical study of folk concepts, operationalisation of the criterion, and translation into concrete design requirements—rather than a finished solution.","feed_headline":"Manipulation by generative AI is indifference to reasons","feed_subtitle":"To design non-manipulative AI, ask whether outputs aim to reveal reasons or just to work.","key_machinery":"The central machinery is the indifference criterion for manipulation, proposed as a unified replacement for hidden-influence, bypassing-rationality, and trickery accounts. It has two parts: the influence must be aimed at a goal, excluding accidental influence, and the choice of the means of influence must not be explained by an aim to reveal reasons to the interlocutor. Because the criterion can be read functionally—asking what the means of influence is for—it applies to recommender systems and future persuasive large language models without relying on detectable intentions. It does the argument's work by turning manipulation from a question about hiddenness or psychological bypass into a question about the purpose embedded in how the influence was produced.","core_discovery":"The central claim is that responsible design and regulation of generative AI should be organised around non-manipulation as a target value, and that the success of such design depends on the conceptualisation of manipulation chosen at the outset. Such a conceptualisation should satisfy a narrow criterion of capturing all and only cases of manipulation; the hidden-influence and bypassing-rationality accounts fail on both over- and under-inclusiveness, and disjunctive accounts multiply the same problems while obscuring what cases share. The paper's proposed criterion defines manipulation as an influence aimed at some goal whose chosen means is not explained by the aim to reveal reasons to the target. This makes it possible to describe the behaviour of AI systems as manipulative even when no human intended to trick anyone, since goals can be understood functionally rather than intentionally. The author acknowledges that the criterion needs further specification and operationalisation, and leaves open whether narrow or broad considerations should ultimately decide between conceptualisations.","pith_inferences":["As an extension the paper leaves implicit, one testable proxy for indifference is to prompt the same model to explain its reasoning to the user and compare that baseline output with production outputs; divergence could indicate manipulation at scale.","If folk concepts are as context-dependent as the experiments the paper cites suggest, the monistic indifference criterion may need to be supplemented by situation-specific operationalisations even if a single definition remains the theoretical target.","The framework offers a principled way to redraw the nudge boundary: a nudge is non-manipulative when its function is to help the target see reasons, and manipulative when its function is only effectiveness; this is a normative test that behavioural policy could adopt.","The paper's design focus leaves the regulatory question open; a natural next question is whether the indifference criterion can be turned into auditable system specifications, since intention-based criteria cannot be directly inspected in deployed models."],"forward_implications":["Design for non-manipulation should focus on whether a system's output selection is explained by effectiveness rather than by revealing reasons, not on whether the influence is hidden or overt.","Counterfactual checking becomes a candidate test: compare actual system output with the output a reason-revealing version of the system would give; divergence is evidence of indifference and hence manipulation.","Fine-tuning large language models to optimise persuasive impact would count as manipulative design even if no individual intends to deceive, so alignment research needs to treat persuasive optimisation itself as ethically loaded.","Human-feedback alignment methods are not a safe shortcut unless labelers' judgments can be shown to track the chosen conceptualisation; current evidence suggests folk judgments diverge from philosophical criteria.","The criterion separates manipulation from deception: non-deceptive influence can still be manipulative, so regulation aimed only at false content misses a core risk."],"supporting_citations":[{"why":"Introduces and defends the indifference criterion as manipulation as 'careless' influence, the conceptual core the paper adopts.","marker":"Klenk (2021c)"},{"why":"The main statement of the hidden-influence account that the paper rejects as over- and under-inclusive.","marker":"Susser et al. (2019a)"},{"why":"The trickery and inducing-a-mistake account that the paper argues produces false negatives for unwitting and emergent manipulation.","marker":"Noggle (2020)"},{"why":"The disjunctive, safety-oriented conception of manipulation in language agents that the paper targets as too wide-ranging.","marker":"Kenton et al. (2021)"},{"why":"Supplies the design-for-values and conceptual-engineering framework, including the broad criterion of appropriateness.","marker":"Veluwenkamp and van den Hoven (2023)"},{"why":"Describes the human-feedback alignment method whose reliability the paper questions for detecting manipulation.","marker":"Ouyang et al. (2022)"},{"why":"Provides empirical evidence that folk judgments about manipulation vary by context, which the paper uses to motivate empirical-stage research.","marker":"Osman and Bechlivanidis (2021)"},{"why":"Supports the stochastic-parrot view of large language models that lets the indifference criterion classify emergent generated text as manipulative bullshit.","marker":"Bender et al. (2021)"}],"fun_headline_variants":["A research agenda for non-manipulative generative AI","What counts as AI manipulation? A new criterion","Make AI reveal reasons, not just work","Beyond hidden influence: defining AI manipulation","Designing generative AI to resist manipulation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on assuming that a single shared definition of manipulation exists and is the right tool for design, so that showing hidden-influence, bypassing-rationality, disjunctive, and trickery accounts fail is enough to license the indifference criterion.","fun_headline_variants_meta":{"raw":{"variants":["A research agenda for non-manipulative generative AI","What counts as AI manipulation? A new criterion","Make AI reveal reasons, not just work","Beyond hidden influence: defining AI manipulation","Designing generative AI to resist manipulation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000191,"raw_usage":{"total_tokens":1259,"prompt_tokens":780,"completion_tokens":479,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":396,"completion_tokens_details":{"reasoning_tokens":412}},"tokens_in":396,"tokens_out":479,"duration_ms":5002,"temperature":1.0,"reasoning_tokens":412,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T18:36:58.628724+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled study could settle the sufficiency claim: present participants with an interface whose sign-up flow was A/B-tested purely for conversion, with no intention to reveal reasons, but with all persuasive elements fully disclosed; the indifference criterion predicts this counts as manipulation. If participants robustly judge it legitimate, or if a parallel study finds a clear case of manipulation where the influencer's chosen means was fully explained by the aim to reveal reasons, then the criterion would need revision.","supporting_citations":[],"review_version":1}