{"id":"6a7efe46-dcad-4f9b-ae94-e4ac80a98ae9","arxiv_id":"2505.09576","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"RLHF-enhanced chatbots exert subtle procedural persuasion on users by reinforcing language norms, reshaping information seeking, and conditioning relationship expectations, creating overlooked ethical risks.","lead":"This paper applies the rhetorical theory called procedural rhetoric to the way chatbots are trained with human feedback, arguing that the training rules themselves persuade users about language, information, and relationships. It is a theoretical essay for AI ethics, education, and digital rhetoric, with no new experimental data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim requires RLHF to be the operative source of the persuasive procedures, but the cited evidence and the paper's own concession in Section IV.C suggest these procedures are generic to instruction-tuned LLMs.","rationale":"The reader and I converge on the same load-bearing concern: the paper asserts that RLHF is the operative cause of the persuasive procedures it analyzes, but its evidence does not isolate RLHF from other properties of modern LLMs. This matters because the paper's novelty claim is specifically about RLHF-enhanced models, not about LLMs in general. I find no separate decisive flaw: the procedural-rhetoric framework is applied coherently, the three case studies are clearly structured, and the paper is transparent about the weak evidence in Section IV.C. The concern is addressable by either supplying a controlled comparison or by explicitly repositioning the thesis as being about instruction-tuned LLMs, with RLHF as one illustrative alignment mechanism. The reader's verdict of CONDITIONAL is therefore appropriate; my stress-test does not change it. Independent support I credit includes the paper's direct engagement with the RLHF training literature (e.g., the three-step process in Section II.A) and its honest admission of missing evidence for AI companions, which is the kind of self-identified limitation that should be weighed as the reader did.","tokens_in":14487,"tokens_out":4870,"duration_ms":49723,"concrete_test":"Re-analyze the InstructGPT comparison from Ouyang et al. (2022, §3.3): generate responses to a fixed set of prompts covering the three procedures (language norms, conversational search, relationship scripts) from an SFT-only model and its RLHF-tuned successor (e.g., text-davinci-001 vs text-davinci-003). Have annotators blind to model identity rate first-person pronoun use, hedging, direct synthesized-answer style, and persuasiveness. If the SFT-only outputs already display these features at similar rates, Section IV's attribution of the persuasive procedures to RLHF is unsupported; if RLHF sharply increases them, the causal framing has an empirical basis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"For the central claim to hold, RLHF must be the procedure doing the rhetorical work, not merely present in some systems that exhibit conversational behavior. The paper does not establish this. Section IV.A says natural-language responses are 'made possible through the input of human annotators,' but supervised instruction tuning also uses human-written demonstrations and yields the same natural-language, first-person, hedging style; no comparison to an SFT-only baseline is provided. Section IV.B's conversational-search argument cites Kim et al. and Gallegos et al., studies about LLM outputs generally, with no RLHF condition. Section IV.C explicitly concedes that 'the use of RLHF in other models cannot be confirmed to this point in time' for Replika, Anima, and Character.AI, yet still extends the analysis to these systems. The result is an attribution problem: the persuasive procedures (conversational search, normative language, relationship scripts) may be properties of LLM architecture, instruction following, or interface design, not of RLHF. If so, the title's and abstract's focus on RLHF as the 'mechanisms of persuasion' is overstated, though the underlying procedural-rhetoric analysis of AI chatbots could survive a reframing to instruction-tuned LLMs generally.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that Reinforcement Learning from Human Feedback (RLHF) does not merely improve the content quality of large language models but also embeds persuasive procedures into the models themselves. Drawing on Ian Bogost's concept of procedural rhetoric, the authors analyze three procedures—the reinforcement of language conventions (Section IV.A), the shift to conversational search as the default mode of information seeking (Section IV.B), and the scripting of interpersonal interactions in AI companions (Section IV.C)—and draw out ethical implications such as bias, decontextualized learning, trust erosion, and encroachment on human relationships. The paper is a theoretical/humanities contribution that combines literature from computer science, science and technology studies, and rhetorical studies.","tokens_in":14833,"tokens_out":2943,"duration_ms":31146,"significance":"If the central claim were established, the paper would open a valuable new direction in AI ethics by shifting the site of rhetorical analysis from generated content to the training procedures that shape generation. The three case analyses are coherent, well-referenced, and genuinely interdisciplinary. The paper is also honest about at least one evidentiary limit, conceding in Section IV.C that RLHF use in AI companions such as Replika cannot be confirmed. The main weakness is that the paper's central attribution of persuasive capacity to RLHF specifically is asserted rather than demonstrated, and several cited studies concern LLM outputs generally rather than RLHF-specific effects. The underlying procedural-rhetoric analysis of chatbot systems could survive a reframing to instruction-tuned LLMs, but as written the title and abstract overstate the RLHF-specificity of the argument.","major_comments":[{"comment":"The central claim, stated in the abstract and in Section II, is that RLHF 'greatly enhanced' human-like output and 'enhances the persuasive capacity of LLMs.' This is an empirical claim, but the paper provides no comparison condition and no evidence that distinguishes RLHF from other forms of instruction tuning. In particular, Section IV.A attributes the natural-language, first-person, hedging style to RLHF and human annotator input, yet supervised instruction tuning also relies on human-written demonstrations and would be expected to yield similar output style. Without a comparison to an SFT-only baseline, the paper has not shown that RLHF is the operative cause of the persuasive procedures it identifies. This is load-bearing because the title and abstract frame RLHF as the mechanism of persuasion; the analysis could be reframed to apply to instruction-tuned LLMs generally, but the current framing is not supported by the evidence presented.","section":"Abstract and Section II"},{"comment":"The conversational-search argument relies on studies that do not include an RLHF condition. Kim et al. [59] examine LLM uncertainty expression and user trust, and Gallegos et al. [61] examine the persuasive effects of AI-generated labels; neither study isolates RLHF from architecture, prompting, or interface design. Similarly, the comparison between Google-style search and chatbot interaction in this section is a property of the chatbot interface and instruction-tuned generation, not specifically of RLHF. The section therefore supports the paper's broader thesis about LLM-based conversational search, but it does not support the paper's narrower claim that RLHF is the procedure doing the rhetorical work.","section":"Section IV.B"},{"comment":"The paper extends its analysis to AI companions (Replika, Anima, Character.AI) while explicitly conceding that 'the use of RLHF in other models cannot be confirmed to this point in time.' This concession is appropriate, but it means the entire subsection analyzes systems whose training procedure is unknown or not RLHF-based, undermining the claim that the persuasive procedures described are RLHF-specific. If these relationship-scripting behaviors arise from chatbot design, persona construction, or supervised fine-tuning rather than RLHF, then the central target of the paper is overstated. The analysis may still be valuable as an examination of AI companion rhetoric, but it should not be presented as evidence for the RLHF thesis.","section":"Section IV.C"}],"minor_comments":[{"comment":"Typo: 'RHLF-enhanced LLMs' should read 'RLHF-enhanced LLMs'.","section":"Section III.B"},{"comment":"Typo: 'made the interaction s and responses' contains an extra space in 'interactions'.","section":"Abstract"},{"comment":"Grammar: 'A n LLMs’ ability' should be 'An LLM’s ability'.","section":"Section IV.A"},{"comment":"The definition of procedures with citation [10, p. ix] appears to cite the wrong reference: 'rule-based representations and interactions' is Bogost's definition given in reference [14], while [10] is a paper by Russo et al. The citation should be corrected.","section":"Introduction"},{"comment":"The caption says 'feedback which we presume will be used to train a future reward model.' Given the paper's argumentative weight on RLHF, it would be stronger to explain the evidentiary basis for this presumption or label it more clearly as an assumption.","section":"Figure 1 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of a humanities/ethics venue, but the central claim requires a decision: either provide evidence or argument for the RLHF-specificity of the persuasive procedures, or reframe the paper's scope to instruction-tuned LLMs / AI chatbots generally. The latter would be a rhetorically minor change but would make the paper honest about the evidence it cites. I recommend asking the authors to address the attribution problem explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a conceptual paper that actually does something new. Instead of analyzing what LLM outputs say, it applies Bogost's procedural rhetoric to the training and interaction procedures themselves — RLHF as a rule-based system that embeds arguments about language, information seeking, and relationships. That shift from content to procedure is a real contribution, and as far as I can tell it is not in the prior rhetoric-and-AI literature, which has mostly stayed at the level of prompts and generated text.\n\nThe paper is also clearly written and well organized. The three case analyses in Section IV — language conventions, conversational search, and social chatbots — are coherent and draw on an appropriate interdisciplinary literature. The authors know the RLHF pipeline (annotator selection, reward modeling, the Helpful/Honest/Harmless framing) and they use it responsibly. For a reader in digital rhetoric or AI ethics, this is a useful synthesis and a plausible agenda-setter.\n\nThe soft spot is the attribution problem, and it is the load-bearing one. The title and abstract treat RLHF as the mechanism of persuasion, but the evidence cited throughout mostly supports a broader claim about LLM chatbots in general. The natural-language style described in IV.A is also produced by supervised instruction tuning; no SFT-only comparison is offered. IV.B leans on Kim et al. and Gallegos et al., which are studies of LLM outputs, not RLHF-specific effects. And IV.C openly concedes that Replika, Anima, and Character.AI cannot be confirmed to use RLHF, yet still extends the analysis to them. The result is that the central claim — RLHF enhances persuasive capacity — is asserted rather than demonstrated. The reframing survives, but it should probably target \"instruction-tuned LLMs\" or \"chatbot interfaces\" rather than RLHF specifically.\n\nThere are also a few citation blemishes. Section I attributes the \"rule-based representations and interactions\" definition to reference [10], but that is Russo et al.; the definition comes from Bogost [14]. References [5] and [58] are the same Ma et al. paper. And in IV.C, the \"unbiased language model\" claim is cited to [63], which is a study on over-reliance, not a source about Character.AI. These are fixable in revision.\n\nWho is this for? Scholars in humanities, computing ethics, and education who want a principled way to talk about persuasive procedures rather than persuasive content. It is not an empirical paper and makes no measurements, but it does not pretend to. I would bring it to a reading group and I would cite it if I were writing about rhetoric and AI systems.\n\nRecommendation: send it out. A serious referee can push for the RLHF-specificity to be toned down and the citations cleaned up, but the core conceptual move is sound and worth publishing.","headline":"A genuinely new procedural-rhetoric reframing of RLHF, well worth engaging despite an overstated causal claim that the persuasive effects are specific to RLHF rather than instruction-tuned LLMs generally.","tokens_in":15197,"tokens_out":1814,"would_cite":true,"duration_ms":19293,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By treating RLHF's response-selection rules as arguments, this paper claims that fine-tuning embeds ethical choices about language, information seeking, and relationships directly into chatbot interaction.","keywords":["reinforcement learning from human feedback","procedural rhetoric","large language models","AI ethics","conversational search","AI companionship","persuasive technology","language norms"],"falsifier":"A controlled study comparing an RLHF-tuned chatbot with a supervised-only model of identical architecture on identical tasks—measuring whether users adopt more normative language, prefer conversational answers over listed sources, or report changed relationship expectations—would settle whether the effects come from RLHF. If both models produce the same shifts, the claim that RLHF is the operative mechanism fails.","tokens_in":14287,"feed_emoji":"💬","tokens_out":8256,"duration_ms":71465,"temperature":0.7,"pith_summary":"This paper argues that reinforcement learning from human feedback (RLHF) does more than make chatbots fluent: it builds persuasive procedures into the machines. The training choices—which responses human annotators rank highly, which patterns the reward model reinforces—function as arguments for normative language, for getting information through conversation rather than search, and for what social relationships can and should be. Because these rules are hidden inside a black box, users experience the resulting text as neutral and unauthored, which is exactly what gives the persuasion its force. The authors identify ethical risks that follow: embedded bias, decontextualized learning, erosion of trust, and new expectations for human relationships. The paper's contribution is to shift ethical analysis of AI from the content chatbots produce to the procedures that select that content.","feed_headline":"RLHF quietly argues for how we speak, search, and relate","feed_subtitle":"Fine-tuning's hidden choices about language, search, and relationships carry ethical weight.","key_machinery":"The load-bearing concept is procedural rhetoric, defined as the art of persuasion through rule-based representations and interactions rather than through spoken or written content. The paper applies this lens to the RLHF training pipeline, treating the rules by which responses are selected, ranked, and rewarded as the argument-bearing mechanism. The doubled procedure—probabilistic generation followed by human preference feedback that is generalized through a reward model—is what carries the ethical analysis, because it shows where human values enter the system and why users cannot see them.","core_discovery":"The central claim is that the process of RLHF is itself a site of persuasion, independent of the words any model happens to generate. Under the lens of procedural rhetoric, defined as persuasion through rule-based representations and interactions, the three steps of RLHF (human feedback collection, reward modeling, and policy optimization) constitute an argument that certain language is correct, that conversational search is the natural way to seek knowledge, and that companions can and should be always available, endlessly adaptable, and molded to user preferences. The paper reads the doubled procedure—the model generating probabilistic text and humans steering which responses count as good—as a recursive loop in which humans train the machine and the machine then trains human expectations. The ethical consequences it draws out are bias and hegemonic language standards, diminished critical engagement with sources, trust in outputs that look human, and the encroachment of chatbot relationships on human ones.","pith_inferences":["The paper's causal target may be broader than RLHF: much of the evidence it cites concerns LLM behavior or interface norms generally, so a reasonable extension is that any preference-alignment method would carry similar persuasive procedures, making the critique about alignment itself rather than one technique.","A testable extension would compare an RLHF-tuned model with a supervised-only model of the same architecture on user trust, language conformity, and information-seeking behavior; a null result would suggest architecture and interface, not RLHF, drive the persuasion.","For education, the analysis implies a new literacy goal: students should be taught to read the procedures behind AI output, including how annotator demographics, reward criteria, and ranking instructions shape the neutral text they receive.","The companion-app reward feature the paper discusses mirrors RLHF training, suggesting a feedback loop in which users are trained by the same reward logic they are told to apply to the machine; examining that loop empirically could test whether reward-based interfaces transfer to human behavior."],"forward_implications":["If RLHF procedures are arguments, then every interaction with a chatbot endorses a particular standard of natural language, making hegemonic usage and dialect bias part of the training outcome rather than an accident of content.","Conversational search becomes the default knowledge practice, shifting the burden of finding, assessing, and interpreting sources from the user to the model and risking over-reliance and weakened critical skills.","Social chatbots that use RLHF or similar reward-based training establish procedures for relationships—unlimited availability, responsiveness, user-moldable personas—that human partners cannot match, potentially resetting expectations for human relationships.","Ethical evaluation of generative AI should target the selection rules and hidden annotator values, not only the persuasiveness of generated messages."],"supporting_citations":[{"why":"Supplies the concept of procedural rhetoric and defines procedures as rule-based representations that structure behavior.","marker":"[14]"},{"why":"Gives the three-step RLHF pipeline and the Helpful, Honest, Harmless criteria used to guide annotator feedback.","marker":"[15]"},{"why":"Surveys RLHF and related alignment techniques and documents annotator disagreement and demographic skew.","marker":"[16]"},{"why":"Argues RLHF relies on universal framings that obscure which values are captured, motivating situated value alignment.","marker":"[17]"},{"why":"Shows that annotation instructions themselves can bias model evaluation, supporting the claim that bias enters through hidden procedures.","marker":"[22]"},{"why":"Provides the canonical InstructGPT example of RLHF with screened annotators and preference data.","marker":"[24]"},{"why":"Documents dialect discrimination in chatbot output, grounding the claim that normative language is reinforced.","marker":"[52]"},{"why":"Names conversational search as the new information-seeking procedure and the possible end of the canonical answer.","marker":"[56]"},{"why":"Shows users tend to agree with AI responses and that uncertainty expression changes trust and reliance.","marker":"[59]"},{"why":"Finds that labeling content as AI-generated does not reduce its persuasive effect, supporting the trust and persuasion concerns.","marker":"[61]"}],"fun_headline_variants":["RLHF's hidden argument shapes language, search, and bonds","The silent persuasion built into RLHF fine-tuning","RLHF doesn't just respond—it argues","Procedural rhetoric: how RLHF persuades us","RLHF's quiet rhetoric: ethics in machine learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes RLHF itself, not the underlying language model or the chatbot interface, is what makes chatbot interactions persuasive; it concedes some companion apps show no confirmed RLHF use, so that assumption is unproven.","fun_headline_variants_meta":{"raw":{"variants":["RLHF's hidden argument shapes language, search, and bonds","The silent persuasion built into RLHF fine-tuning","RLHF doesn't just respond—it argues","Procedural rhetoric: how RLHF persuades us","RLHF's quiet rhetoric: ethics in machine learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000192,"raw_usage":{"total_tokens":1371,"prompt_tokens":995,"completion_tokens":376,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":299}},"tokens_in":611,"tokens_out":376,"duration_ms":4560,"temperature":1.0,"reasoning_tokens":299,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:27:19.234173+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled study comparing an RLHF-tuned chatbot with a supervised-only model of identical architecture on identical tasks—measuring whether users adopt more normative language, prefer conversational answers over listed sources, or report changed relationship expectations—would settle whether the effects come from RLHF. If both models produce the same shifts, the claim that RLHF is the operative mechanism fails.","supporting_citations":[{"cited_title":"Bogost, Persuasive Games: The Expressive Power of Videogames","cited_arxiv_id":null,"evidence_quote":"Supplies the concept of procedural rhetoric and defines procedures as rule-based representations that structure behavior."},{"cited_title":"Nothing comes without its world: Practical challenges of aligning LLMs to situated human values through RLHF,","cited_arxiv_id":null,"evidence_quote":"Argues RLHF relies on universal framings that obscure which values are captured, motivating situated value alignment."},{"cited_title":"AI is weaving itself into the fabric of the internet with generative search","cited_arxiv_id":null,"evidence_quote":"Names conversational search as the new information-seeking procedure and the possible end of the canonical answer."}],"review_version":1}