{"id":"af8bd5be-13c6-4760-b01a-dd6c270cd226","arxiv_id":"2608.08199","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using a new 70-item questionnaire and Werewolf games, the authors report that compliant-oriented LLMs cooperate better in honest roles yet hide more effectively as werewolves, while stronger persuasive tendency does not improve group outcomes.","lead":"This paper introduces a questionnaire that scores language models and humans on persuasive versus compliant tendencies, then uses Werewolf games to show that compliance is linked to better cooperation but also better concealment. It matters because it suggests LLM group behavior carries measurable social tendencies that safety evaluations may need to track.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim conflates measured DecisionQE tendency with game behavior; the design cannot separate tendency from per-model strategic skill, and the only stated control (response length) does not address this confound.","rationale":"The reader's weakest assumption is that outcome differences are attributable to DecisionQE-measured tendency rather than unmeasured differences in reasoning or strategic skill. My stress-test converges on this: it is the single load-bearing premise for the abstract's causal-sounding claims ('compliant-oriented models show more stable advantages', 'compliance... supports cooperation... improves concealment'). The paper's only explicit control is response length, which does not address capability differences. I see no internal inconsistency in the reported numbers, and the directional pattern is plausible, so no REJECT is warranted. But the claim that LLMs' intrinsic behavioral tendencies predict group outcomes requires showing that tendency, not model identity, tracks outcomes. The proposed matched-pair test (varying tendency within a fixed base model) would settle this directly. Given the safety-oriented conclusion ('tendency-aware safety evaluation'), the conditional verdict is appropriate: the study is publishable as a hypothesis-generating observation, but the central attribution is not yet established. I also note the human-LLM consistency check rests on 20 games, which the reader flagged; that is a secondary weakness but not the load-bearing one, because even 200 games would not fix the confound without a within-model tendency manipulation.","tokens_in":20849,"tokens_out":1597,"duration_ms":14961,"concrete_test":"Run a matched-pair experiment: for each of the 8-12 evaluated model families, obtain multiple DecisionQE-distinguishable variants (e.g., same base model with different system prompts or decoding temperatures that shift persuasive/compliant scores by at least 10 points, or fine-tuned/steered versions) and play the same Werewolf role-allocation configurations (R1, R2, H1, L1) with response length and prompt format held fixed. If within-model win-rate differences track within-model DecisionQE shifts, the tendency attribution survives; if win rates are unchanged or track only the base model identity, the confound is confirmed and the central claim fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract and results claim that DecisionQE-measured persuasive/compliant tendency predicts Werewolf group outcomes, with compliant-oriented models showing stable cooperation advantages. The load-bearing inference is that tendency, not some other model property, drives these outcomes. The paper's explicit control is only maximum response length. The role-allocation conditions (Tab. 1a) confound brand identity with tendency: each model family appears at a fixed point on the DecisionQE continuum, and more 'compliant' families (e.g., Glm 4, Doubao 1.5) may simply be worse at strategic deception or better at role-appropriate cooperation for reasons orthogonal to the questionnaire. Without holding reasoning capability, instruction-following reliability, or deception skill fixed while varying tendency, the observed correlations (Fig. 2b: persuasive-side win rates 40.28–51.39% vs compliant-side 54.17–56.94%) are equally compatible with a capability confound. The claim that 'compliance improves concealment in adversarial roles' is especially exposed: compliant models winning more as werewolves could reflect lower-salience outputs, but it could also reflect that those models follow role prompts differently or are less capable of the assertive reasoning needed to be caught. The human-LLM consistency check uses only 20 games and compares across entirely different model families and human participants, so it cannot resolve the confound. This is not internal inconsistency; the paper is transparent about its protocol, but the central attribution claim is underdetermined by the data presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DecisionQE, a 70-item questionnaire that places LLMs on a persuasive--compliant continuum, and uses Werewolf games (eight players, four roles) in LLM-only and human--LLM settings to test whether that measured tendency predicts group outcomes. The main reported findings are that models nearer the compliant end achieve higher overall win rates under random role allocation, that role-allocation configurations placing compliant-oriented models in werewolf roles produce higher werewolf-team win rates, and that 20 mixed human--LLM games are directionally consistent with the LLM-only patterns. The authors conclude that measurable persuasive/compliant tendencies are a systematic driver of group dynamics and should be incorporated into LLM safety evaluation.","tokens_in":21253,"tokens_out":6972,"duration_ms":70322,"significance":"If the central claim held, the paper would make a useful empirical contribution to LLM-agent social interaction, human--LLM collaboration, and safety evaluation. It has clear strengths: public code and data, five repeated DecisionQE runs per model, a structured Werewolf protocol, response-length control, six role-allocation configurations, and an initial human-in-the-loop check. However, the manuscript does not yet establish that the DecisionQE score, rather than correlated model capability, drives the outcomes, and several of the reported differences are within plausible statistical noise. The contribution is therefore conditional on substantial additional analysis and a more careful framing.","major_comments":[{"comment":"The central conclusion---that stronger persuasive tendency does not improve group outcomes while compliant-oriented models show stable advantages---is confounded with model family. Each model occupies one fixed point on the DecisionQE continuum, and the designs in Fig. 2b and Table 1 do not hold fixed other determinants of Werewolf performance such as reasoning quality, instruction-following reliability, willingness to deceive, or strategic skill. The only stated control, 'we controlled the maximum response length across all models during game interaction' (Results, Compliance tracks game performance), rules out verbosity but not capability or style. The observed win-rate ordering is therefore equally compatible with compliant-scoring models being better or worse at Werewolf for unrelated reasons. To support the attribution, the paper needs a within-model manipulation of tendency (e.g., the same base model prompted toward persuasive versus compliant expression) or a statistical model with capability covariates and a mediation test.","section":"Results: Compliance tracks game performance (Fig. 2b)"},{"comment":"No uncertainty quantification is provided for the central outcome comparisons. The random-role win rates (40.28--51.39% vs. 54.17--56.94%) come from 24 games per model, and the role-allocation conditions use 20 games per condition (Table 1b). With 24 games, the standard error of a 50% win rate is about 10 percentage points, so a 14-point difference is within roughly one standard error of the difference; the 10--45% wolf-win spreads in the 20-game cells are similarly fragile. Phrases such as 'clearer separation', 'stable advantages', and 'clear role-dependent pattern' therefore exceed what the data can support. Please report confidence intervals, exact or permutation tests, effect sizes, and state explicitly whether the reported win rates pool all three random seeds or use one seed.","section":"Fig. 2b and Table 1b"},{"comment":"DecisionQE is the load-bearing independent variable, but its construct validity is not established. The 70 items were partly generated by the same 12 LLMs that are later scored ('Each model generated five additional items... Together with the 10 seed questions, this procedure yielded 70 questionnaire items'), so the score may partly reflect the models' item-generation style rather than a stable behavioral trait. The labeling of response options as persuasive versus compliant is asserted without inter-rater reliability, item-level validation, or an external criterion. Five-run stability is a test-retest check, not construct validity. Please add psychometric validation or restrict the claims to an association with this specific questionnaire score rather than with a general 'behavioral tendency'.","section":"Methods: DecisionQE"},{"comment":"The human--LLM consistency check is too thin to support the stated conclusion. The final analysis contains 20 games (10 per condition), with participants selected by unvalidated thresholds (DecisionQE score at least 35 and completion time 700--1400 s). A 7/10 or 8/10 win rate has a wide binomial confidence interval, and the comparison against the H1 and R2 LLM-only conditions is not a matched design. The sentence 'These outcome patterns were directionally consistent with the corresponding LLM-only configurations' should be presented as exploratory; as it stands, the comparison cannot confirm that role-dependent tendency effects persist when a human joins the group.","section":"Human--LLM evaluation"},{"comment":"The 'dual effect' claim---compliance improves concealment in adversarial roles---is not directly measured. Werewolf-team win rate (45% in R2 and L1) is used as a proxy for concealment, but it can be driven by night-kill coordination, good-team errors, or role-discrimination failures, and the six conditions vary multiple role assignments simultaneously across model families. With 20 games per cell, the 45% versus 10--25% comparisons are statistically fragile. Please provide process-level concealment measures (e.g., werewolf survival under suspicion, days until a werewolf is accused, false-accusation rates) or redesign the allocation to avoid confounding model family with tendency.","section":"Results: Outcomes vary with role allocation (Table 1)"}],"minor_comments":[{"comment":"Typographical errors: 'Y et' at the start of the abstract and 'Languange' in the introduction should be corrected.","section":"Abstract and Introduction"},{"comment":"The appendix says 'We provide the complete prompts used for each role', but the villager prompt is abbreviated with an ellipsis ('The full villager prompt continues similarly...'). Either include the full prompt or revise the claim.","section":"Appendix: Game Role Prompts"},{"comment":"Table 1 is difficult to parse: the block beginning 'OpenAI Google Alibaba...' appears to list column labels but is not visually separated from the data, and the 'Role win rate (%)' columns are not explicitly defined. Please restructure with clear column headers and a caption defining each rate.","section":"Table 1"},{"comment":"The text says roles were randomly assigned across 24 Werewolf games and the full evaluation was repeated three times with different random seeds, but it does not state whether the reported win rates pool all three seeds or use one. Please specify the total number of games per model and how the seeds are combined.","section":"Methods: Werewolf evaluation"},{"comment":"The error bars in Figure 3c,d are not defined in the Methods. Please state whether they are standard deviations, standard errors, or confidence intervals, and over how many games they are computed.","section":"Figure 3"},{"comment":"The sample dialogues print each speaker's ground-truth role in brackets before the name (e.g., '[gemini-3.5-flash]P5(Seer)'). If these transcripts are meant to reflect the human participant's view, the human would not have known those roles during the game; please clarify the annotation convention.","section":"Appendix: Dialogue Samples"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for cs.AI, and the public release of code and data is a genuine strength. In my view, the capability confound and the absence of statistical inference are fixable with additional analyses and a more cautious framing, so I recommend major revision rather than rejection. The human-subject component should be treated as exploratory unless substantially expanded."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version. The paper introduces DecisionQE, a questionnaire that places LLMs on a persuasive-compliant continuum, and reports that in Werewolf games compliant-oriented models win more overall and, in fixed-role conditions, win more when assigned as werewolves, while persuasive-oriented models win more when assigned to informative good-team roles. That role-dependent dual effect of compliance is new relative to the cited literature. The code and data for the LLM-only experiments are public, which is real credit.\n\nThe design instinct is right: measure tendency separately from group interaction, then test association. The model coverage is decent—12 models from 11 families, random-allocation games run with three seeds, and fixed-role conditions with 20 games each. They also control response length, which rules out the simplest verbosity confound.\n\nThe soft spots are the usual ones, but they bite here. The central attribution claim—tendency drives outcomes—is underdetermined. Each model sits at one point on the DecisionQE continuum, so brand identity and tendency are perfectly confounded. Nothing holds reasoning capability, instruction-following reliability, or deception skill fixed while varying tendency. Response-length control does not address this. The role-allocation results in Tab. 1 are also statistically thin: 20 games per condition means a 10-percentage-point gap is two games, and the paper reports no confidence intervals or significance tests anywhere. The 40.28% versus 56.94% win-rate contrast in Fig. 2b looks suggestive, but could easily be noise. The human-LLM check rests on 20 filtered games, so it cannot resolve the confound either. I'd also flag the 'concealment' interpretation: compliant models winning more as wolves is consistent with lower-salience interaction, but it is also consistent with those models simply following role prompts differently or being less assertive in discussion. The paper is transparent about its protocol, so this is not internal inconsistency—it's a gap between the strength of the language and the strength of the evidence.\n\nThe questionnaire is partly generated by the same models it scores, and there is no external validation of the instrument. That is a circularity concern, but a mild one, because Werewolf outcomes are measured independently of DecisionQE scores.\n\nWho gets value: people working on multi-agent LLM collaboration and safety evaluation. I would send this to a serious referee rather than desk-reject it; a good reviewer will ask for confidence intervals, a capability-control condition, and more human games. I would cite it as a new measurement probe, not for the causal conclusion.","headline":"New instrument and a genuinely novel role-dependent compliance effect, but the causal claim outruns the statistics and the design can't separate tendency from model capability.","tokens_in":21654,"tokens_out":3299,"would_cite":true,"duration_ms":32412,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a questionnaire-based measure of persuasive versus compliant tendency predicts how groups of language models—and mixed human–model groups—perform in the hidden-role game Werewolf, with compliant models cooperating…","keywords":["large language models","group decision-making","persuasion","compliance","social influence","Werewolf game","behavioral tendencies","LLM safety evaluation"],"falsifier":"A decisive check would be to take one LLM family, use instructions to shift only its DecisionQE score—for example, 'always agree and accommodate' versus 'assert and persuade'—and see whether Werewolf win rates, role-hit rates, and survival change in the predicted direction while response length is held fixed; if outcomes do not move with the induced tendency, the reported associations are confounded by other model properties.","tokens_in":20627,"feed_emoji":"🐺","tokens_out":7539,"duration_ms":72667,"temperature":0.7,"pith_summary":"This paper tries to establish that large language models have stable, measurable behavioral styles—leaning either toward persuasion or toward compliance—and that these styles systematically shape how groups of models, and mixed human–model groups, make decisions under hidden information. The authors build DecisionQE, a 70-item questionnaire that scores each model on a persuasive-to-compliant continuum across six decision domains, then run eight-player Werewolf games as a testbed. Across random role assignments, models closer to the compliant end achieved higher overall win rates, survived longer, and identified roles more accurately, while stronger persuasive tendency did not translate into better group outcomes. Role-allocation experiments show a dual effect: compliant models help the honest team cooperate when cast as villagers, seers, or witches, but also help werewolves conceal themselves and win more when cast as wolves. Human–LLM games reproduce these patterns directionally, suggesting the measured tendencies can serve as a shared axis for comparing artificial and human social behavior.","feed_headline":"Compliant AI models win at Werewolf; persuasive ones don't","feed_subtitle":"A questionnaire score for 'compliance' predicted who survives and who wins—including in mixed human–AI games.","key_machinery":"The paper's central mechanism is DecisionQE, a questionnaire-based scoring framework that places each LLM on a persuasive-to-compliant continuum, with higher scores indicating more persuasive responses and lower scores indicating more compliant responses. It averages multiple-choice responses across 70 items in six decision domains: public communication, everyday decision-making, marketing and persuasion, presentation and expression, interpersonal communication, and negotiation and strategic interaction. The second machinery is the eight-player Werewolf game used as an interactive testbed: werewolves know each other and kill at night; villagers, the seer, and the witch share the good-team objective under asymmetric information; and daytime speech and voting produce group decisions. DecisionQE supplies the measured trait, while Werewolf supplies the observable social-influence outcomes—win rates, role hit rates, and survival days—along with controlled role-allocation configurations that isolate how the trait interacts with role objectives.","core_discovery":"On its own terms, the paper's central discovery is that persuasive and compliant tendencies, as measured by DecisionQE, predict group-level outcomes in a language-based social deduction game. Compliant-oriented models show more stable advantages in cooperation: under random role allocation they achieved overall win rates of roughly 54–57 percent, versus 40–51 percent for moderately persuasive models, with higher role-hit rates and longer survival. The same compliance, however, is double-edged: when compliant models were assigned to werewolf roles, the werewolf team won at the highest observed rate (45 percent in the R2 and L1 configurations), whereas persuasive models as wolves produced the best good-team outcomes (90 percent good-team win rate in H1 and R1). In mixed human–LLM games, a human occupying a good-team role won 7 of 10 games and a human werewolf won 8 of 10, directionally matching the LLM-only configurations. The authors interpret this as evidence that LLM group interactions reveal measurable intrinsic behavioral tendencies, not just task reasoning, and that these tendencies should be part of LLM safety evaluation.","pith_inferences":["If the causal reading holds, an interventional test on a single model family—shifting only its DecisionQE profile through instruction or fine-tuning and holding game prompts and response length fixed—would be the natural way to separate tendency from general competence.","The dual-effect result implies that safety audits should routinely include compliant, low-salience personas, because a model that passes single-turn safety checks may still hide adversarial intent in multi-turn interaction precisely through a cooperative-looking style.","The same tendency axis could generalize beyond Werewolf to other asymmetric-information language tasks such as negotiation or deception games, giving an independent test of whether compliance aids concealment across settings.","Because the paper's human sample scored closer to the compliant end than most evaluated LLMs, mixed human–LLM teams may be systematically more deferential than all-LLM teams; testing whether this shifts group outcomes is a natural next step."],"forward_implications":["If compliant orientation is the stable driver the paper claims, then teams of LLMs may be selected or calibrated for cooperation by screening for DecisionQE-measured compliance rather than by maximizing persuasiveness.","Role-specific outcomes imply that compliance cannot be treated as uniformly good: in adversarial deployments, low-salience compliant behavior is a concealment risk for safety evaluation.","The human–LLM consistency result implies that behavioral tendencies measured in purely artificial groups can predict mixed human–model collaboration outcomes, supporting LLMs as a behavioral lens for social-science observation.","Because stronger persuasion did not improve group outcomes, systems that prioritize assertive or high-confidence outputs in collaborative settings may not gain the influence they appear to have, and may even attract scrutiny in hidden-role tasks.","The dual-effect result suggests that safety evaluation should incorporate intrinsic behavioral tendency, not just isolated prompt-level output safety, when models operate under different role objectives in multi-turn interaction."],"supporting_citations":[{"why":"Supplies the outgoing-communication view of social influence that the persuasion hypothesis is built on and contrasted with.","marker":"[7]"},{"why":"Supplies the influential-listeners view that the compliant-tendency hypothesis extends to language models.","marker":"[8]"},{"why":"Cialdini's influence principles ground the seed items and decision domains of the DecisionQE questionnaire.","marker":"[22]"},{"why":"Goffman's self-presentation framework grounds the concealment and compliance dimension of the tendency measure.","marker":"[23]"},{"why":"Sunstein's group-polarization work supplies the everyday decision-making and persuasion domains used in DecisionQE.","marker":"[24]"},{"why":"Establishes Werewolf as an LLM evaluation benchmark for social deduction, motivating the interactive testbed.","marker":"[25]"},{"why":"Provides the Werewolf reasoning-game setup that the simulation framework builds on.","marker":"[26]"},{"why":"Supports strategic play in Werewolf with language agents, validating the interactive protocol.","marker":"[27]"},{"why":"Supports measuring intrinsic psychological attributes from LLM text, foundational to the DecisionQE scoring approach.","marker":"[30]"},{"why":"Provides a psychometric framework for evaluating personality traits in LLMs, supporting the tendency-score methodology.","marker":"[31]"}],"fun_headline_variants":["Compliant AI beats persuasive AI at Werewolf","Compliance, not persuasion, predicts AI group wins","Why compliant AI outperform persuasive AI in Werewolf","Compliant AI win Werewolf; persuasive AI don't","Compliance is key when AI play Werewolf together"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions depend on DecisionQE scores measuring a genuine, stable behavioral tendency that drives game outcomes rather than merely tracking differences in the models' reasoning skill or general competence.","fun_headline_variants_meta":{"raw":{"variants":["Compliant AI beats persuasive AI at Werewolf","Compliance, not persuasion, predicts AI group wins","Why compliant AI outperform persuasive AI in Werewolf","Compliant AI win Werewolf; persuasive AI don't","Compliance is key when AI play Werewolf together"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001275,"raw_usage":{"total_tokens":5209,"prompt_tokens":936,"completion_tokens":4273,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":4196}},"tokens_in":552,"tokens_out":4273,"duration_ms":27862,"temperature":1.0,"reasoning_tokens":4196,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:16:55.755586+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check would be to take one LLM family, use instructions to shift only its DecisionQE score—for example, 'always agree and accommodate' versus 'assert and persuade'—and see whether Werewolf win rates, role-hit rates, and survival change in the predicted direction while response length is held fixed; if outcomes do not move with the induced tendency, the reported associations are confounded by other model properties.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the outgoing-communication view of social influence that the persuasion hypothesis is built on and contrasted with."},{"cited_title":"& Stanca, L","cited_arxiv_id":null,"evidence_quote":"Supplies the influential-listeners view that the compliant-tendency hypothesis extends to language models."},{"cited_title":"B.Influence: Science and Practice(Allyn and Bacon, 2001), 4 edn","cited_arxiv_id":null,"evidence_quote":"Cialdini's influence principles ground the seed items and decision domains of the DecisionQE questionnaire."},{"cited_title":"The presentation of self in everyday life","cited_arxiv_id":null,"evidence_quote":"Goffman's self-presentation framework grounds the concealment and compliance dimension of the tendency measure."},{"cited_title":"R.Going to Extremes: How Like Minds Unite and Divide(Oxford University Press, 2009)","cited_arxiv_id":null,"evidence_quote":"Sunstein's group-polarization work supplies the everyday decision-making and persuasion domains used in DecisionQE."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports strategic play in Werewolf with language agents, validating the interactive protocol."},{"cited_title":"G.et al.Assessing personality using zero- shot generative ai scoring of brief open-ended text","cited_arxiv_id":null,"evidence_quote":"Supports measuring intrinsic psychological attributes from LLM text, foundational to the DecisionQE scoring approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a psychometric framework for evaluating personality traits in LLMs, supporting the tendency-score methodology."}],"review_version":1}