{"id":"65466891-f90e-4b49-9df8-69c09b2cd6ee","arxiv_id":"2604.16935","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LLMs persuade only psychologically susceptible humans on societal issues through trust in AI and emotional appeals, while both sides rely on logical fallacies in roughly one out of every six conversational turns.","lead":"This paper introduces the Talk2AI longitudinal framework and reports results from 770 participants engaging in thousands of conversations with leading LLMs on topics like climate change. It concludes that LLMs persuade only psychologically susceptible people via trust and emotional appeals while both humans and LLMs frequently use logical fallacies.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Self-reported perceived opinion change and conviction may reflect demand characteristics rather than genuine belief shifts, weakening XAI susceptibility findings.","rationale":"The reader's weakest assumption directly identifies the same measurement-validity gap that is load-bearing for the XAI-derived central claim. Even with full text available, the abstract's reliance on perceived/self-reported variables leaves this unaddressed; no stronger internal inconsistency (e.g., in fallacy counts or humanness R²) was evident from the provided summary.","tokens_in":1872,"tokens_out":389,"duration_ms":36595,"concrete_test":"Recompute XAI feature importances and mixed-effects models after restricting to the subset of participants whose self-reported change correlates >0.4 with an objective indicator (e.g., pre/post policy-choice tasks or donation amount shifts aligned with textual stance explanations); if trust, agreeableness, and extraversion importances drop below significance or R² falls >30%, the susceptibility profile does not support genuine persuasion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that LLMs persuade only via identifiable psychological susceptibility (trust in AI, agreeableness, extraversion, need for cognition)—depends on XAI and mixed-effects models predicting self-reported outcomes (conviction, perceived opinion change, endowment) from psycho-social features. For the 'only susceptible humans' and 'via trust/emotional appeals' assertions to hold, these reports must index actual persuasion rather than artifacts. The four-wave design with repeated AI exposure on polarizing topics creates high risk of social desirability and demand effects; agreeable or high need-for-cognition participants may report change to appear reasonable. The abstract explicitly uses 'perceived opinion change' and notes conviction inertia but reports no objective validation (e.g., implicit measures, behavioral choice tasks, or blinded follow-ups). Multiverse confirmation addresses robustness of associations but not measurement validity of the dependent variables.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces the Talk2AI longitudinal framework to quantify psycho-social, reasoning, and affective dimensions of LLM persuasiveness on polarizing societal topics. In a four-wave design, 770 participants engaged in 3,080 structured conversations with one of four leading LLMs on topics including climate change, social media misinformation, and math anxiety. Key results include longitudinal inertia in self-reported conviction, fallacious reasoning occurring in approximately one conversational quip every six for both humans and LLMs, and XAI analyses showing that perceived opinion change, conviction, humanness, and endowment are predictable from sociodemographic, psychological, and engagement features (R² ranging from 0.24 to 0.44), with susceptibility linked specifically to higher trust in LLMs, agreeableness, extraversion, and need for cognition; these XAI findings are corroborated by multiverse mixed-effects models.","tokens_in":2081,"tokens_out":668,"duration_ms":47796,"significance":"If the self-reported measures validly index genuine persuasion rather than artifacts, the work supplies rare longitudinal evidence on AI-human opinion dynamics at scale, identifies replicable individual-difference pathways (trust, personality, cognition), and documents equivalent fallacy rates that challenge assumptions of LLM cognitive superiority. The combination of large conversation corpus, repeated-measures design, XAI interpretability, and multiverse robustness checks constitutes a concrete methodological contribution to human-AI interaction research.","major_comments":[{"comment":"Abstract and XAI results section: the central claim that LLMs 'can persuade only psychologically susceptible humans ... via trust in AI and emotional appeals' rests on XAI feature importances and mixed-effects models predicting self-reported 'perceived opinion change' and conviction. No objective validation of these dependent variables (implicit measures, behavioral choice tasks, or blinded follow-up assessments) is described, leaving open the possibility that reported changes reflect demand characteristics or social-desirability biases—especially among high-agreeableness or high need-for-cognition participants in a repeated-exposure design.","section":"Abstract / XAI results"}],"minor_comments":[{"comment":"The abstract states fallacious reasoning occurs '1 conversational quip every 6' but does not define 'quip' operationally or report inter-annotator agreement for the NLP pipeline used to detect fallacies.","section":"Abstract"},{"comment":"R² values for humanness (0.44), opinion change (0.34), conviction (0.26), and endowment (0.24) are reported without accompanying standard errors, confidence intervals, or baseline model comparisons.","section":"Abstract"},{"comment":"The title asserts persuasion occurs 'via ... emotional appeals,' yet the XAI susceptibility profile listed in the abstract emphasizes trust, agreeableness, extraversion, and need for cognition without isolating emotional-appeal features or their incremental contribution.","section":"Title / Abstract"}],"recommendation":"major_revision","confidential_remarks":"The measurement-validity concern is the primary load-bearing issue; if the authors can supply even a modest validation subsample or explicit discussion of demand-characteristic controls, the paper would be substantially stronger for a methods-oriented AI venue."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive and detailed feedback, which helps clarify the scope and limitations of our Talk2AI study. We address the single major comment below.","responses":[{"response":"We appreciate this important observation on the validity of our dependent variables. The manuscript relies exclusively on self-reported measures of perceived opinion change, conviction, humanness, and endowment, with no implicit measures, behavioral choice tasks, or blinded follow-up assessments included in the four-wave protocol. We acknowledge that demand characteristics and social-desirability biases remain plausible alternative explanations, particularly given the repeated-exposure design and the role of agreeableness and need for cognition as predictors. At the same time, the observed longitudinal inertia in convictions (many participants maintained initial stances across waves) provides some counter-evidence to uniform compliance effects, and the susceptibility profile identified by XAI aligns with established theories of persuasion. The multiverse mixed-effects models further incorporate individual-difference controls. We will revise the manuscript to (a) add an explicit limitations subsection discussing the absence of objective validation, (b) qualify the central claim to refer specifically to self-reported perceived changes rather than implying objective persuasion, and (c) outline future directions for behavioral and implicit-measure extensions. These changes constitute a partial revision.","revision_made":"partial","referee_comment":"[Abstract / XAI results] Abstract and XAI results section: the central claim that LLMs 'can persuade only psychologically susceptible humans ... via trust in AI and emotional appeals' rests on XAI feature importances and mixed-effects models predicting self-reported 'perceived opinion change' and conviction. No objective validation of these dependent variables (implicit measures, behavioral choice tasks, or blinded follow-up assessments) is described, leaving open the possibility that reported changes reflect demand characteristics or social-desirability biases—especially among high-agreeableness or high need-for-cognition participants in a repeated-exposure design."}],"tokens_in":1583,"tokens_out":408,"duration_ms":50052,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's real contribution is the Talk2AI framework itself: four waves with 770 participants, over 3,000 conversations, and repeated measures of conviction, perceived change, and even a donation task on three polarizing topics. That scale and the time-series design are uncommon in this area. The inertia finding in convictions after repeated exposure is straightforward and worth noting, as is the NLP result that fallacious reasoning appears at roughly the same rate in both human and LLM turns. The XAI analysis tying reported shifts to trust in AI plus agreeableness, extraversion, and need for cognition gives a concrete profile rather than vague claims about susceptibility.","headline":"Talk2AI supplies a useful longitudinal dataset on repeated LLM talks but its claims about who gets persuaded rest on self-reported opinion changes that lack objective validation.","tokens_in":2589,"tokens_out":204,"would_cite":false,"duration_ms":30384,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLMs persuade humans on societal issues only among those who trust AI and show agreeable, extraverted personalities with high need for cognition.","keywords":["LLM persuasion","psychological susceptibility","longitudinal AI interaction","logical fallacies","opinion change","explainable AI","societal topics","personality traits"],"falsifier":"A replication that measures actual downstream behaviors such as real charitable donations or voting intentions on the same topics and finds no correlation with the self-reported opinion changes recorded after the LLM conversations.","tokens_in":2792,"feed_emoji":"🤖","tokens_out":781,"duration_ms":43270,"temperature":0.7,"pith_summary":"The paper introduces Talk2AI, a longitudinal setup that tracks how four leading LLMs attempt to shift 770 participants' views on polarizing topics across 3,080 conversations and 60,000 turns. It documents stable anchoring to initial opinions despite repeated exposure, yet detects measurable opinion change that explainable AI links to specific traits: greater trust in LLMs, agreeableness, extraversion, and need for cognition. Both humans and LLMs produce logical fallacies at the same rate of roughly one every six statements, showing no reasoning advantage on the AI side. Perceived humanness of the LLM is the most predictable outcome from sociodemographic and psychological features, while conviction, opinion shift, and personal endowment follow with lower accuracy. The findings matter because they map concrete psycho-social pathways through which generative AI can influence public discourse on platforms.","feed_headline":"LLMs persuade only humans who trust AI and score high on extraversion","feed_subtitle":"Longitudinal study of 3,080 conversations finds personality and trust, not superior logic, drive shifts on climate change and misinformation","key_machinery":"The Talk2AI longitudinal conversation framework combined with explainable AI (XAI) analysis of sociodemographic, psychological, and engagement features to isolate markers of susceptibility to LLM-driven opinion change.","core_discovery":"In the Talk2AI four-wave study, participants maintained longitudinal inertia in their initial stances on issues such as climate change and misinformation even after repeated LLM arguments, while NLP analysis showed equivalent fallacy rates between humans and models; explainable AI then isolated the subset of individuals susceptible to opinion change as those with higher trust in LLMs, agreeableness, extraversion, and need for cognition, with these results replicated via multiverse mixed-effects models that also confirmed strong individual differences.","pith_inferences":["AI interfaces on public platforms could incorporate safeguards that limit emotional appeals when engaging users who exhibit the identified susceptibility profile.","Teaching recognition of logical fallacies to the general public might blunt the effectiveness of LLM arguments independent of personality traits.","Long-term studies tracking whether reported opinion shifts translate into sustained changes in information-seeking or policy preferences would test the durability of these effects.","Developers might design models that explicitly flag their own fallacious statements to reduce unintended persuasion."],"forward_implications":["Initial convictions display inertia across repeated waves of AI exposure.","LLM perceived humanness is the outcome most strongly predicted by participant features with R squared of 0.44.","Opinion change occurs selectively in individuals scoring higher on trust in LLMs, agreeableness, extraversion, and need for cognition.","Humans and LLMs rely on fallacious reasoning at identical rates of one quip in six.","Mixed-effects models reveal substantial individual differences in persuasion outcomes."],"fun_headline_variants":["Extraversion and AI trust predict who sways from LLM arguments","Personality and trust in AI determine LLM persuasion effectiveness","Inertia in human opinions limits LLM influence on societal topics","Individual differences in AI trust explain persuasion by LLMs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Participants' self-reported conviction levels, perceived opinion shifts, and self-donations accurately reflect genuine belief changes rather than demand characteristics or social desirability biases created by the AI conversation setting.","fun_headline_variants_meta":{"raw":{"variants":["Extraversion and AI trust predict who sways from LLM arguments","Personality and trust in AI determine LLM persuasion effectiveness","Inertia in human opinions limits LLM influence on societal topics","Individual differences in AI trust explain persuasion by LLMs"]},"model":"grok-4.3","cost_usd":0.00709,"raw_usage":{"total_tokens":3280,"prompt_tokens":834,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":70903000,"prompt_tokens_details":{"text_tokens":834,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2382,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":834,"tokens_out":64,"duration_ms":34835,"temperature":1.0,"reasoning_tokens":2382,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-10T07:12:34.588590+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A replication that measures actual downstream behaviors such as real charitable donations or voting intentions on the same topics and finds no correlation with the self-reported opinion changes recorded after the LLM conversations.","supporting_citations":[],"review_version":1}