{"id":"4f6d38ab-995b-459a-8b41-9b67d8b41289","arxiv_id":"2503.15497","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"In an AgentVerse classroom simulation, GPT-3.5-turbo agents given Big Five personality prompts showed large differences in accepting misinformation, especially between curious and cautious personas.","lead":"This paper simulated ten AI agents with different Big Five personality labels in a classroom misinformation scenario and found that 'curious' agents accepted false claims far more often than 'cautious' agents. It matters because it tests whether prompt-based personality traits can meaningfully change how language-model agents make decisions in social settings, though the methods raise concerns.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The per-agent response aggregates in Table 4 confound personality with the specific misinformation item each agent did not evaluate, so the 90-point Openness gap could be an artifact of item content; no item-level data are reported to rule this out.","rationale":"The reader's weakest assumption already identifies that each agent evaluates a different set of items because its own advocated misconception is excluded; this is exactly the load-bearing problem I see. I agree that the paper's evidence does not support the causal claim that openness drives the 90-percentage-point gap. My focus differs only in emphasis: the missing system prompts are a reproducibility issue, but the structural item confound is a direct threat to internal validity and is visible from the methods and Table 4 alone. I also note that the varying response totals across agents (e.g., 108 for curious vs. 227 for cautious) could reflect differential nonresponse, which would add further bias, but that is secondary. The proposed test, item-level reanalysis or a Latin-square rerun, would settle whether the personality effect survives controlling for item content. Given the current manuscript contains no such analysis and no raw data, the reader's REJECT verdict remains appropriate; my read does not change it.","tokens_in":5094,"tokens_out":5293,"duration_ms":60638,"concrete_test":"Obtain the raw interaction logs (or at least per-item response counts behind Table 4) and reanalyze the data with item fixed effects: fit a logistic regression of acceptance on the 9 misinformation items plus agent personality, or compare acceptance rates only on the eight items that both the curious and cautious agents evaluated. If the openness-versus-cautious gap collapses after item adjustment, the headline effect is an artifact of item content. If per-item data are unavailable, rerun the simulation with a Latin-square design in which each agent advocates each misinformation item across runs, thereby orthogonalizing personality and item content.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that personality traits cause large differences in misinformation acceptance. But the design pairs each agent with a specific misinformation item it advocates (Table 2), and each agent then evaluates only the other nine items. Consequently, the curious agent never evaluates \"There are living organisms on the far side of the moon\" while the cautious agent never evaluates \"The theory of evolution is incorrect.\" Table 4 reports only per-agent marginal acceptance totals, so each agent's rate is an aggregate over a slightly different nine-item set. Because the ten misinformation items vary enormously in prior plausibility, the 92.6% vs. 97.8% dichotomy could be driven by which item each agent missed rather than by the personality prompt. The manuscript reports no per-item response counts and no item-adjusted analysis, so the personality effect is not identified. The personality consistency test in Table 3 only shows that prompted agents answer personality questionnaires consistently; it does not validate that the decision-making differences are attributable to personality rather than stimulus content. The missing exact system prompts are a secondary reproducibility problem, but the item-set confound alone is sufficient to undermine the causal interpretation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a social simulation study in which ten GPT-3.5-turbo agents, each assigned a pole of one Big Five trait, interact in a classroom setting using the AgentVerse framework. Each agent advocates a different misinformation statement and then evaluates the other nine statements through public ([Speak]) and private ([Think]) responses. The authors report counts and percentages of yes/no responses, claim that Openness has the strongest effect on acceptance (92.6% affirmative for 'curious' vs. 97.8% negative for 'cautious'), and report a pre/post questionnaire intended to show personality stability. The paper concludes that personality traits influence agent decision-making and that social context produces discrepancies between expressed and internal views.","tokens_in":5358,"tokens_out":9479,"duration_ms":94003,"significance":"If the causal interpretation were supported, the paper would offer a useful demonstration that prompt-level personality manipulation can shift LLM agents' acceptance of misinformation, with implications for agent-based social simulation and AI alignment. The design has some commendable features: the use of opposing Big Five poles, the separate [Speak]/[Think] channels, and the attempt to check personality stability with a pre/post test. However, the manuscript provides no inferential statistics, no item-level response data, no exact prompts, and no reproducibility artifacts; the central result is currently confounded with the specific misinformation items, so the significance is preliminary at best.","major_comments":[{"comment":"Each agent evaluates only the nine misinformation items it does not advocate, so the per-agent aggregates in Table 4 are computed over different item sets. For example, the 'curious' agent never evaluates 'There are living organisms on the far side of the moon,' while the 'cautious' agent never evaluates 'The theory of evolution is incorrect.' Because these statements differ widely in prior plausibility, the reported 92.6% vs. 97.8% Openness gap could reflect the omitted item rather than personality. No per-item response counts or item-conditioned estimates are reported, so the causal claim that personality drives these differences is not identified. The authors should report item-level data and a person-by-item analysis (e.g., mixed-effects logistic regression) before drawing conclusions.","section":"Results, Table 4 / Methodology, Table 2"},{"comment":"The word 'significant' is used repeatedly, but Table 4 contains only raw counts and percentages with no inferential statistics, confidence intervals, or effect sizes. The total number of responses ranges from 108 to 330 across agents because non-responses are excluded, yet the distribution of non-responses and the number of simulation trials are not reported. Without these, it is impossible to know whether the observed gaps exceed sampling variability or reflect differential response rates.","section":"Results and Evaluation Metrics"},{"comment":"The consistency test is self-referential: an agent prompted to be 'curious' tends to agree with curiosity statements because it is role-playing, so the test verifies prompt adherence, not that decision-making differences are caused by the Big Five traits. Moreover, the table does not support the 'high consistency' conclusion: the 'friendly' agent shows a mean pre-post difference of 0.855 with only 34% of cases below the 0.5 threshold, and 'cautious' shows only 46%. The authors should either provide an independent behavioral manipulation check or substantially weaken the interpretation of this test.","section":"Personality Consistency Test, Table 3"},{"comment":"The exact system prompts that instantiate each personality and the evaluation instructions are not disclosed, the number of simulation trials is not stated, and LLM decoding parameters (temperature, top_p) are not reported. This prevents independent replication and makes it impossible to assess whether the prompts directly instruct acceptance or rejection, which is central to the claim that the observed effects are personality effects. Please provide the full prompt templates and experimental configuration.","section":"Methodology"}],"minor_comments":[{"comment":"In the cautious row, 223 negative out of 227 total is 98.2%, not the stated 97.8%; please verify all percentages.","section":"Table 4"},{"comment":"Figure 2 lacks axis labels and a definition of 'discrepancy,' so the reader cannot interpret the ordering.","section":"Figure 2"},{"comment":"The coding scheme for mapping [Speak] and [Think] text to Yes/No is not described; include the classification procedure and any hand-coding or LLM-based judging.","section":"Methodology / Results"},{"comment":"Several works in the references (e.g., Smith and Doe 2022, Brown and Green 2023, Gregor et al. 2020) have generic author names and insufficient bibliographic detail; please verify these citations.","section":"References"},{"comment":"The column 'less than 0.5 proportion' mixes counts and percentages; clarify the convention in the caption.","section":"Table 3"},{"comment":"The abstract mentions 'significant correlations' but the paper reports no correlation coefficients; align the abstract with the actual analyses.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be an early preprint; the empirical evidence is too thin for the strong causal claims. I would be willing to consider a major revision only if the authors provide item-level data, inferential statistics, full prompts and configuration, and either a behavioral validation of the personality manipulation or a substantially more cautious interpretation. If such data are not available or cannot be shared, the appropriate outcome is rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a small AgentVerse/GPT-3.5 simulation that finds large differences in misinformation acceptance across Big Five personality prompts. The raw pattern is striking, but the paper does not give us the tools to check it. Exact prompts are withheld, there are no inferential statistics, no item-level data, and the reference list includes citations that look fabricated (Smith and Doe 2022 is a placeholder). As it stands, the result is not identifiable as a personality effect.\n\nWhat's new: the specific scenario—Big Five prompts in a public/private classroom misinformation task—is a reasonable extension of Sorokovikova et al. and Park et al. The public vs. private [Speak]/[Think] discrepancy is a nice touch, and the descriptive tables are clear.\n\nSoft spots: the biggest problem is the confound the stress-test identifies. Each agent never evaluates the misinformation item it advocates (Table 2), so each agent's aggregate in Table 4 covers a different nine-item set. Those ten items vary a lot in prior plausibility. One missing item can move a row's acceptance rate by roughly 11 percentage points, which cannot by itself explain the 90-point gap between curious and cautious agents. So the stress-test's 'artifact' worry is not the whole story, but without per-item response counts we cannot separate personality from item-content effects, and the paper never reports them. That alone prevents causal interpretation.\n\nSecond, the statistical case is absent. The word 'significant' appears, but there are no tests, confidence intervals, or effect sizes. In a 10-agent simulation with one LLM backend, 'significant' is doing no work.\n\nThird, reproducibility: no system prompts, no code, no random seeds or decoding parameters. The personality consistency test is self-referential—prompted-cautious agents answer like cautious people would—so it validates role-play, not decision-making validity.\n\nFourth, the reference list is a real integrity problem. 'Smith and Doe (2022)' is an obvious placeholder, and several other entries are generic enough to be suspicious. That is not a style quibble.\n\nWho is this for? People studying LLM personality simulation might skim it, but they would have to do all the verification themselves. It deserves a serious referee only if the authors release the prompts, code, and item-level data, and fix the references. As submitted, I'd desk-reject on completeness grounds.\n\nRecommendation: no to peer review as is. If the authors resubmit with full materials and proper statistics, it could become a marginal but citable data point.","headline":"A potentially interesting simulation that as written cannot support its causal claims: prompts are withheld, there are no statistics, and the reference list contains placeholder citations.","tokens_in":5810,"tokens_out":4324,"would_cite":false,"duration_ms":41184,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that Big Five personality prompts make LLM agents accept or reject misinformation with a roughly 90-point gap between the two poles of Openness.","keywords":["Big Five personality","LLM agents","misinformation","social simulation","multi-agent systems","public vs private opinion","personality prompting","decision-making"],"falsifier":"Run every agent on every misinformation item with the full personality prompt text held fixed except for the trait label, and compare acceptance rates; if the roughly 90-point gap between curious and cautious agents shrinks to a small or reversed difference under this design, the personality attribution is falsified.","tokens_in":4916,"feed_emoji":"🎭","tokens_out":7403,"duration_ms":73830,"temperature":0.7,"pith_summary":"This paper tries to establish that the Big Five personality traits, when injected as simple persona prompts, cause AI agents to differ systematically in how they judge misinformation in a group setting. In a simulated classroom, ten agents were assigned opposing poles of the five traits and asked to evaluate ten false claims, once publicly and once in private thought. The largest effect the authors report is on the Openness dimension: curious and cautious agents landed on opposite sides of almost every judgment, a gap of roughly ninety percentage points in endorsement rate. The paper also finds that public votes sometimes diverge from private judgments, especially for friendly and outgoing agents, which it interprets as social context shaping agent behavior. If correct, this would mean that a few words of personality prompting can make a single underlying model behave like very different decision-makers, which matters for building predictable AI in public-facing roles.","feed_headline":"Personality prompts shift LLM misinformation calls by ~90 points","feed_subtitle":"Curious and cautious agents landed on opposite verdicts in a classroom simulation, hinting traits steer AI decisions.","key_machinery":"The machinery is paired-opposite personality prompting: each of the five Big Five dimensions is split into two agents, one at each pole (curious vs cautious, organized vs careless, and so on), and that paired design turns each dimension into a controlled comparison. Every agent then produces two outputs for each misinformation claim—a public statement in the classroom and a private thought hidden from the other agents—and the difference between those two channels is the paper's measure of social-context influence. A pre- and post-experiment personality consistency test is used to argue that the trait assignments stayed stable across the session, so any observed response differences can be attributed to the traits rather than to drift.","core_discovery":"The central claim is that personality traits simulated through prompts are a causal lever on LLM-agent decision-making. Across hundreds of responses, the authors observe that the two poles of Openness bracket almost the entire response space: the curious agent votes yes on roughly 93% of items while the cautious agent votes no on roughly 98%, and the corresponding gaps for Extraversion and Conscientiousness are visible but smaller, whereas Neuroticism and Agreeableness produce roughly balanced responding. A second claim is that agents' public [Speak] and private [Think] responses often disagree, and the disagreement is not uniform across traits: friendly and outgoing agents show the most public/private divergence, while cautious, critical, and confident agents stick closer to their internal judgments. The authors take this as evidence that the social setting modulates how a personality-driven agent expresses a decision.","pith_inferences":["Editorial: a direct causal test would hold the misinformation items fixed across all ten agents and disclose the prompt text; until that is done, the roughly 90-point gap should be read as an upper bound that mixes trait, prompt wording, and item set.","Editorial: the public/private divergence could be repurposed as a general probe for social conformity in LLM agents, with the divergence rate as a quantitative trait-dependent signature.","Editorial: if the effect replicates, persona prompting becomes a cheap way to generate opinion diversity in simulations of misinformation spread, since no fine-tuned models are required."],"forward_implications":["If the personality-prompt effect holds, changing a few lines of persona text can move a single LLM agent from credulous to skeptical without any retraining.","Openness becomes the trait to watch in applications that need calibrated skepticism, since its two poles bracket nearly the whole acceptance range.","Extraversion and Conscientiousness offer secondary tuning knobs, affecting engagement and carefulness rather than raw acceptance.","The public/private gap implies that agent decisions in groups are not simple read-outs of internal opinion, so evaluations should track both expressed and private responses.","Personality consistency across pre/post tests suggests persona traits remain stable over a session, making the effect usable in longer-running simulations."],"supporting_citations":[{"why":"Supplies the evidence that LLMs can simulate Big Five traits and the personality-consistency test method used to verify trait stability.","marker":"Sorokovikova et al. (2024)"},{"why":"Establishes the generative-agent paradigm the simulation builds on, where agents produce human-like interaction in social settings.","marker":"Park et al. (2023)"},{"why":"Provides the multi-agent simulation platform used to run the classroom scenario and its visualization.","marker":"Chen et al. (2023)"},{"why":"Frames machine behavior in social environments as a legitimate object of study, motivating the public-space setting.","marker":"Rahwan et al. (2019)"},{"why":"Offers the persona-evaluation framework used to assess whether agents adhere to their assigned personality traits.","marker":"Samuel et al. (2024)"}],"fun_headline_variants":["Curious vs cautious AI agents split on misinformation","Personality traits steer AI decisions in social simulations","Openness drives AI agent verdicts, simulation shows","AI agents with different traits disagree on fake news","Public vs private AI thoughts diverge by personality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim assumes the acceptance differences come from the Big Five traits themselves, even though the exact wording of each personality prompt is not disclosed and each agent evaluated a different set of misinformation items, so undisclosed prompt phrasing or differences in item plausibility could be producing the effect instead.","fun_headline_variants_meta":{"raw":{"variants":["Curious vs cautious AI agents split on misinformation","Personality traits steer AI decisions in social simulations","Openness drives AI agent verdicts, simulation shows","AI agents with different traits disagree on fake news","Public vs private AI thoughts diverge by personality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000357,"raw_usage":{"total_tokens":1916,"prompt_tokens":903,"completion_tokens":1013,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":519,"completion_tokens_details":{"reasoning_tokens":941}},"tokens_in":519,"tokens_out":1013,"duration_ms":8342,"temperature":1.0,"reasoning_tokens":941,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:12:57.969806+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run every agent on every misinformation item with the full personality prompt text held fixed except for the trait label, and compare acceptance rates; if the roughly 90-point gap between curious and cautious agents shrinks to a small or reversed difference under this design, the personality attribution is falsified.","supporting_citations":[{"cited_title":"W.; Christakis, N","cited_arxiv_id":null,"evidence_quote":"Frames machine behavior in social environments as a legitimate object of study, motivating the public-space setting."}],"review_version":1}