Pith. sign in

REVIEW 5 major objections 6 minor 47 references

Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models

T0 review · 5 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper claims that a questionnaire-based measure of persuasive versus compliant tendency predicts how groups of language models—and mixed human–model groups—perform in the hidden-role game Werewolf, with compliant models cooperating…

desk verdict New instrument and a genuinely novel role-dependent compliance effect, but the causal claim outruns the statistics and the design can't separate tendency from model capability. read the letter →

arxiv 2608.08199 v1 pith:R2M2HSXE submitted 2026-08-08 cs.AI

classification cs.AI
keywords largelanguagemodelsgroupdecision-makingpersuasioncompliancesocialinfluenceWerewolfgamebehavioraltendenciesLLMsafetyevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that large language models have stable, measurable behavioral styles—leaning either toward persuasion or toward compliance—and that these styles systematically shape how groups of models, and mixed human–model groups, make decisions under hidden information. The authors build DecisionQE, a 70-item questionnaire that scores each model on a persuasive-to-compliant continuum across six decision domains, then run eight-player Werewolf games as a testbed. Across random role assignments, models closer to the compliant end achieved higher overall win rates, survived longer, and identified roles more accurately, while stronger persuasive tendency did not translate into better group outcomes. Role-allocation experiments show a dual effect: compliant models help the honest team cooperate when cast as villagers, seers, or witches, but also help werewolves conceal themselves and win more when cast as wolves. Human–LLM games reproduce these patterns directionally, suggesting the measured tendencies can serve as a shared axis for comparing artificial and human social behavior.

What carries the argument

The paper's central mechanism is DecisionQE, a questionnaire-based scoring framework that places each LLM on a persuasive-to-compliant continuum, with higher scores indicating more persuasive responses and lower scores indicating more compliant responses. It averages multiple-choice responses across 70 items in six decision domains: public communication, everyday decision-making, marketing and persuasion, presentation and expression, interpersonal communication, and negotiation and strategic interaction. The second machinery is the eight-player Werewolf game used as an interactive testbed: werewolves know each other and kill at night; villagers, the seer, and the witch share the good-team objective under asymmetric information; and daytime speech and voting produce group decisions. DecisionQE supplies the measured trait, while Werewolf supplies the observable social-influence outcomes—win rates, role hit rates, and survival days—along with controlled role-allocation configurations that isolate how the trait interacts with role objectives.

What would settle it

A decisive check would be to take one LLM family, use instructions to shift only its DecisionQE score—for example, 'always agree and accommodate' versus 'assert and persuade'—and see whether Werewolf win rates, role-hit rates, and survival change in the predicted direction while response length is held fixed; if outcomes do not move with the induced tendency, the reported associations are confounded by other model properties.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that persuasive and compliant tendencies, as measured by DecisionQE, predict group-level outcomes in a language-based social deduction game. Compliant-oriented models show more stable advantages in cooperation: under random role allocation they achieved overall win rates of roughly 54–57 percent, versus 40–51 percent for moderately persuasive models, with higher role-hit rates and longer survival. The same compliance, however, is double-edged: when compliant models were assigned to werewolf roles, the werewolf team won at the highest observed rate (45 percent in the R2 and L1 configurations), whereas persuasive models as wolves produced the best good-team outcomes (90 percent good-team win rate in H1 and R1). In mixed human–LLM games, a human occupying a good-team role won 7 of 10 games and a human werewolf won 8 of 10, directionally matching the LLM-only configurations. The authors interpret this as evidence that LLM group interactions reveal measurable intrinsic behavioral tendencies, not just task reasoning, and that these tendencies should be part of LLM safety evaluation.

Load-bearing premise

The conclusions depend on DecisionQE scores measuring a genuine, stable behavioral tendency that drives game outcomes rather than merely tracking differences in the models' reasoning skill or general competence.

Editorial extensions

If this is right

  • If compliant orientation is the stable driver the paper claims, then teams of LLMs may be selected or calibrated for cooperation by screening for DecisionQE-measured compliance rather than by maximizing persuasiveness.
  • Role-specific outcomes imply that compliance cannot be treated as uniformly good: in adversarial deployments, low-salience compliant behavior is a concealment risk for safety evaluation.
  • The human–LLM consistency result implies that behavioral tendencies measured in purely artificial groups can predict mixed human–model collaboration outcomes, supporting LLMs as a behavioral lens for social-science observation.
  • Because stronger persuasion did not improve group outcomes, systems that prioritize assertive or high-confidence outputs in collaborative settings may not gain the influence they appear to have, and may even attract scrutiny in hidden-role tasks.
  • The dual-effect result suggests that safety evaluation should incorporate intrinsic behavioral tendency, not just isolated prompt-level output safety, when models operate under different role objectives in multi-turn interaction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the causal reading holds, an interventional test on a single model family—shifting only its DecisionQE profile through instruction or fine-tuning and holding game prompts and response length fixed—would be the natural way to separate tendency from general competence.
  • The dual-effect result implies that safety audits should routinely include compliant, low-salience personas, because a model that passes single-turn safety checks may still hide adversarial intent in multi-turn interaction precisely through a cooperative-looking style.
  • The same tendency axis could generalize beyond Werewolf to other asymmetric-information language tasks such as negotiation or deception games, giving an independent test of whether compliance aids concealment across settings.
  • Because the paper's human sample scored closer to the compliant end than most evaluated LLMs, mixed human–LLM teams may be systematically more deferential than all-LLM teams; testing whether this shifts group outcomes is a natural next step.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper introduces DecisionQE, a 70-item questionnaire that places LLMs on a persuasive--compliant continuum, and uses Werewolf games (eight players, four roles) in LLM-only and human--LLM settings to test whether that measured tendency predicts group outcomes. The main reported findings are that models nearer the compliant end achieve higher overall win rates under random role allocation, that role-allocation configurations placing compliant-oriented models in werewolf roles produce higher werewolf-team win rates, and that 20 mixed human--LLM games are directionally consistent with the LLM-only patterns. The authors conclude that measurable persuasive/compliant tendencies are a systematic driver of group dynamics and should be incorporated into LLM safety evaluation.

Significance. If the central claim held, the paper would make a useful empirical contribution to LLM-agent social interaction, human--LLM collaboration, and safety evaluation. It has clear strengths: public code and data, five repeated DecisionQE runs per model, a structured Werewolf protocol, response-length control, six role-allocation configurations, and an initial human-in-the-loop check. However, the manuscript does not yet establish that the DecisionQE score, rather than correlated model capability, drives the outcomes, and several of the reported differences are within plausible statistical noise. The contribution is therefore conditional on substantial additional analysis and a more careful framing.

major comments (5)
  1. [Results: Compliance tracks game performance (Fig. 2b)] The central conclusion---that stronger persuasive tendency does not improve group outcomes while compliant-oriented models show stable advantages---is confounded with model family. Each model occupies one fixed point on the DecisionQE continuum, and the designs in Fig. 2b and Table 1 do not hold fixed other determinants of Werewolf performance such as reasoning quality, instruction-following reliability, willingness to deceive, or strategic skill. The only stated control, 'we controlled the maximum response length across all models during game interaction' (Results, Compliance tracks game performance), rules out verbosity but not capability or style. The observed win-rate ordering is therefore equally compatible with compliant-scoring models being better or worse at Werewolf for unrelated reasons. To support the attribution, the paper needs a within-model manipulation of tendency (e.g., the same base model prompted toward persuasive versus compliant expression) or a statistical model with capability covariates and a mediation test.
  2. [Fig. 2b and Table 1b] No uncertainty quantification is provided for the central outcome comparisons. The random-role win rates (40.28--51.39% vs. 54.17--56.94%) come from 24 games per model, and the role-allocation conditions use 20 games per condition (Table 1b). With 24 games, the standard error of a 50% win rate is about 10 percentage points, so a 14-point difference is within roughly one standard error of the difference; the 10--45% wolf-win spreads in the 20-game cells are similarly fragile. Phrases such as 'clearer separation', 'stable advantages', and 'clear role-dependent pattern' therefore exceed what the data can support. Please report confidence intervals, exact or permutation tests, effect sizes, and state explicitly whether the reported win rates pool all three random seeds or use one seed.
  3. [Methods: DecisionQE] DecisionQE is the load-bearing independent variable, but its construct validity is not established. The 70 items were partly generated by the same 12 LLMs that are later scored ('Each model generated five additional items... Together with the 10 seed questions, this procedure yielded 70 questionnaire items'), so the score may partly reflect the models' item-generation style rather than a stable behavioral trait. The labeling of response options as persuasive versus compliant is asserted without inter-rater reliability, item-level validation, or an external criterion. Five-run stability is a test-retest check, not construct validity. Please add psychometric validation or restrict the claims to an association with this specific questionnaire score rather than with a general 'behavioral tendency'.
  4. [Human--LLM evaluation] The human--LLM consistency check is too thin to support the stated conclusion. The final analysis contains 20 games (10 per condition), with participants selected by unvalidated thresholds (DecisionQE score at least 35 and completion time 700--1400 s). A 7/10 or 8/10 win rate has a wide binomial confidence interval, and the comparison against the H1 and R2 LLM-only conditions is not a matched design. The sentence 'These outcome patterns were directionally consistent with the corresponding LLM-only configurations' should be presented as exploratory; as it stands, the comparison cannot confirm that role-dependent tendency effects persist when a human joins the group.
  5. [Results: Outcomes vary with role allocation (Table 1)] The 'dual effect' claim---compliance improves concealment in adversarial roles---is not directly measured. Werewolf-team win rate (45% in R2 and L1) is used as a proxy for concealment, but it can be driven by night-kill coordination, good-team errors, or role-discrimination failures, and the six conditions vary multiple role assignments simultaneously across model families. With 20 games per cell, the 45% versus 10--25% comparisons are statistically fragile. Please provide process-level concealment measures (e.g., werewolf survival under suspicion, days until a werewolf is accused, false-accusation rates) or redesign the allocation to avoid confounding model family with tendency.
minor comments (6)
  1. [Abstract and Introduction] Typographical errors: 'Y et' at the start of the abstract and 'Languange' in the introduction should be corrected.
  2. [Appendix: Game Role Prompts] The appendix says 'We provide the complete prompts used for each role', but the villager prompt is abbreviated with an ellipsis ('The full villager prompt continues similarly...'). Either include the full prompt or revise the claim.
  3. [Table 1] Table 1 is difficult to parse: the block beginning 'OpenAI Google Alibaba...' appears to list column labels but is not visually separated from the data, and the 'Role win rate (%)' columns are not explicitly defined. Please restructure with clear column headers and a caption defining each rate.
  4. [Methods: Werewolf evaluation] The text says roles were randomly assigned across 24 Werewolf games and the full evaluation was repeated three times with different random seeds, but it does not state whether the reported win rates pool all three seeds or use one. Please specify the total number of games per model and how the seeds are combined.
  5. [Figure 3] The error bars in Figure 3c,d are not defined in the Methods. Please state whether they are standard deviations, standard errors, or confidence intervals, and over how many games they are computed.
  6. [Appendix: Dialogue Samples] The sample dialogues print each speaker's ground-truth role in brackets before the name (e.g., '[gemini-3.5-flash]P5(Seer)'). If these transcripts are meant to reflect the human participant's view, the human would not have known those roles during the game; please clarify the annotation convention.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the DecisionQE tendency scores and Werewolf outcomes are measured independently, and no fitted parameter or self-citation chain forces the reported associations.

full rationale

The paper's central derivation separates individual tendency measurement from group-level outcome evaluation. DecisionQE scores are obtained from standardized questionnaire responses, while Werewolf outcomes come from an interactive game with independently defined win, survival, and role-hit metrics. The statistical associations reported in Fig. 2 and Table 1 are correlational: model tendency scores are not constructed from game outcomes, nor are game outcomes computed from tendency scores. No equation in the paper defines one quantity in terms of the other, and no parameter is fitted to game data and then renamed as a prediction. The control for response length is a stated experimental control, not a fitted parameter. The only potentially self-referential element is that some questionnaire items were generated by the same model families that were later scored, but this does not make the measured tendencies equivalent to the game results by construction; it is a measurement-construction concern rather than a circular derivation. No load-bearing self-citations appear; the cited persuasion and network literature is background context, not the basis of the empirical claim. Therefore, under the requirement to exhibit a specific reduction, no circular step can be identified, and the appropriate verdict is no significant circularity.

Assumptions & free parameters 2 free parameters · 3 assumptions · 1 invented entities

The central claim rests on the construct validity of DecisionQE, on Werewolf as a proxy for group decision-making, and on the assumption that outcome differences reflect tendency rather than model competence. No parameters are fitted to derive the results, but two hand-chosen human eligibility thresholds affect the human sub-study.

free parameters (2)
  • Human DecisionQE inclusion cutoff = 35 (score)
    Participants with DecisionQE score below 35 are excluded from the human-LLM games; this hand-chosen threshold shapes the human cohort and all reported human results.
  • Human questionnaire completion time window = 700-1400 seconds
    Eligibility requires finishing DecisionQE within 700 to 1400 seconds; this hand-chosen window may select for faster or more attentive participants and affects generalizability.
assumptions (3)
  • domain assumption DecisionQE scores reflect stable, intrinsic behavioral tendencies of LLMs
    The paper shows within-model stability across five runs, but provides no external validation against established persuasion or personality scales; the construct validity of the single continuum is assumed in Methods, DecisionQE.
  • domain assumption Werewolf game performance is a valid proxy for language-mediated group decision-making and social influence
    The paper motivates Werewolf as capturing asymmetric information and collective voting, but treats game outcomes as evidence of real-world social influence without demonstrating transfer to non-game settings.
  • ad hoc to paper The authors' labeling of questionnaire options as persuasive versus compliant is valid and consistent across all 70 items
    No inter-rater reliability or independent validation is reported; all LLMs cluster near the persuasive end of the scale, suggesting the labels or item design may carry an implicit bias.
invented entities (1)
  • Unified persuasive-compliant tendency continuum (DecisionQE score)
    purpose: To quantify each LLM's or human's orientation toward assertion versus accommodation, and to link that orientation to group outcomes in Werewolf.
    The score is defined solely by the authors' option labels and model-generated items. No external psychometric validation is provided, and the paper's own Werewolf results are the only evidence offered for its predictive value.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models." pith.science (2026). https://pith.science/paper/R2M2HSXE

@misc{pith2026260808199,
  author       = {Pith},
  title        = {Pith review of: Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/R2M2HSXE}},
  note         = {Machine review of arXiv:2608.08199}
}
read the original abstract

Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and humans. Yet it remains unclear whether their influence is driven by persuasion-oriented expression or compliance-oriented accommodation. We introduce DecisionQE, a questionnaire-based framework for measuring each model's persuasive and compliant tendencies across multiple decision scenarios, and use the Werewolf game as an interactive testbed to study their effects on social influence and group outcomes under asymmetric information. Across experiments, stronger persuasive tendency does not significantly improve group outcomes, whereas compliant-oriented models show more stable advantages in cooperation. We further reveal a dual effect of compliance: it supports cooperation in honest roles but improves concealment in adversarial roles. These findings suggest that LLM group interactions reveal not only task outcomes, but also measurable patterns of intrinsic behavioral tendency. LLMs can therefore serve as a lens for sociological observation of language-mediated interaction, while highlighting the need to incorporate behavioral tendencies into safety evaluation of LLM systems.

Figures

Figures reproduced from arXiv: 2608.08199 by the authors.

Figure 1
Figure 1. Evaluation framework. a. DecisionQE Construction, which assesses persuasive and compliant tendencies across six decision-related domains: public communication, everyday decision-making, marketing and persuasion, presentation and expression, interpersonal communication, and negotiation and strategic interaction. b. Werewolf Pipeline. Candidate LLMs are assigned to four roles: werewolves, villagers, seer, and witch. W… view at source ↗
Figure 2
Figure 2. LLM tendency profiles and Werewolf outcomes. a. DecisionQE Score Distribution. DecisionQE tendency-score distributions across the evaluated models; higher scores indicate more persuasive responses and lower scores indicate more compliant responses. b. Werewolf Game Win Rates. Overall model, good-team and werewolf-team win rates for representative models grouped by tendency category. c. Survival and Role Identificati… view at source ↗
Figure 3
Figure 3. Human tendency profiles and outcomes in mixed human–LLM games. a. Human Cohort Selection. Of 20 recruited participants, 12 met the eligibility criteria for the Werewolf evaluation, in which one human interacted with seven LLMs. b. Human and LLM Score Comparison. DecisionQE domain-score distributions for human participants and LLMs. c. Good Replacement Survival Days. Mean survival rounds in games with a human good-te… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Web-based interface for the human–LLMs Werewolf game. The interface displays the player list, visible role information, event and speech history, and phase-specific decision panel. In this example, the human participant is assigned as a villager, while the identities o…
Figure 5
Figure 5. Figure 5: Web-based interface for the daytime speech phase of the human–LLMs Werewolf game. The interface displays the current speaking order, previous system events, and speeches from other players, while providing an input box for the human participant to submit their own spee…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 35 canonical work pages

  1. [1]

    J., Holmes, K

    Flusberg, S. J., Holmes, K. J., Thibodeau, P. H., Nabi, R. L. & Matlock, T. The psychology of framing: How everyday language shapes the way we think, feel, and act.Psychol. Sci. Public Interest25, 105–161, DOI: 10.1177/15291006241246966 (2024)

  2. [2]

    P.et al.An inclusive, real-world investigation of persuasion in language and verbal behavior.J

    Ta, V . P.et al.An inclusive, real-world investigation of persuasion in language and verbal behavior.J. Comput. Soc. Sci.5, 883–903 (2022)

  3. [3]

    & Hewstone, M

    Tausch, N., Schmid, K. & Hewstone, M. The social psychologyof intergroup relations. InHandbook on peace education, 75–86 (Psychology Press, 2011)

  4. [4]

    Ostrom, E.Governing the Commons: The Evolu- tion of Institutions for Collective Action(Cambridge University Press, 1990)

  5. [5]

    Arrow, K. J. A difficulty in the concept of social welfare.J. Polit. Econ.58, 328–346 (1950). 8/26

  6. [6]

    SUNSTEIN, C. R. The law of group polarization.J. Polit. Philos.10, 175–195 (2002)

  7. [7]

    Brandts, J., Giritligil, A. E. & Weber, R. A. An exper- imental study of persuasion bias and social influence in networks.Eur. Econ. Rev.80, 214–229 (2015)

  8. [8]

    & Stanca, L

    Corazzini, L., Pavesi, F., Petrovich, B. & Stanca, L. Influential listeners: An experiment on persuasion bias in social networks.Eur. Econ. Rev.56, 1276– 1288 (2012)

Show all 47 references
  1. [9]

    InConference on Neural Information Processing Systems, 1877–1901 (2020)

    Brown, T.et al.Language models are few-shot learn- ers. InConference on Neural Information Processing Systems, 1877–1901 (2020)

  2. [10]

    Xiao, C.et al.Densing law of llms.Nat. Mach. Intell. 1–11 (2025)

  3. [11]

    Guo, D.et al.Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature645, 633–640, DOI: 10.1038/s41586-025-09422-z (2025)

  4. [12]

    A., Tihanyi, N

    Ferrag, M. A., Tihanyi, N. & Debbah, M. From llm reasoning to autonomous ai agents: A comprehensive review.IEEE Access(2026)

  5. [13]

    Guo, D.et al.Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature645, 633–638 (2025)

  6. [14]

    Truhn, D., Reis-Filho, J. S. & Kather, J. N. Large lan- guage models should be used as scientific reasoning engines, not knowledge databases.Nat. Medicine29, 2983–2984 (2023)

  7. [15]

    InACM SIGKDD Conference on Knowl- edge Discovery and Data Mining, 4314–4325 (2024)

    Zhang, W.et al.A multimodal foundation agent for financial trading: Tool-augmented, diversified, and generalist. InACM SIGKDD Conference on Knowl- edge Discovery and Data Mining, 4314–4325 (2024)

  8. [16]

    InAAAI Conference on Artificial Intelligence(2024)

    Zhang, J.et al.Exploring collaboration mechanisms for llm agents: A social psychology view. InAAAI Conference on Artificial Intelligence(2024)

  9. [17]

    Wang, W.et al.A survey of llm-based agents in medicine: How far are we from baymax? InAn- nual Meeting of the Association for Computational Linguistics(2025)

  10. [18]

    Why artificial intelligence needs sociol- ogy of knowledge: parts i and ii.AI & society40, 1249–1263 (2025)

    Collins, H. Why artificial intelligence needs sociol- ogy of knowledge: parts i and ii.AI & society40, 1249–1263 (2025)

  11. [19]

    InAAAI Conference on Artificial Intelligence, vol

    Chang, S.et al.Llms generate structurally realistic social networks but overestimate political homophily. InAAAI Conference on Artificial Intelligence, vol. 19, 341–371 (2025)

  12. [20]

    & Ferrara, E

    Jiang, J. & Ferrara, E. Social-llm: Modeling user behavior at scale using language models and social network data.Sci7, 138 (2025)

  13. [21]

    & Tang, X

    Zheng, W. & Tang, X. Simulating social network with llm agents: an analysis of information propagation and echo chambers. InInternational Symposium on Knowledge and Systems Sciences, 63–77 (2024)

  14. [22]

    B.Influence: Science and Practice(Allyn and Bacon, 2001), 4 edn

    Cialdini, R. B.Influence: Science and Practice(Allyn and Bacon, 2001), 4 edn

  15. [23]

    The presentation of self in everyday life

    Goffman, E. The presentation of self in everyday life. InSocial theory re-wired, 450–459 (Routledge, 2023)

  16. [24]

    R.Going to Extremes: How Like Minds Unite and Divide(Oxford University Press, 2009)

    Sunstein, C. R.Going to Extremes: How Like Minds Unite and Divide(Oxford University Press, 2009)

  17. [25]

    & Chen, F

    Bailis, S., Friedhoff, J. & Chen, F. Werewolf arena: A case study in llm evaluation via social deduction (2024). ArXiv preprint arXiv:2407.13943

  18. [26]

    ArXiv preprint arXiv:2402.02330

    Wu, S.et al.Enhance reasoning for large language models in the game werewolf (2024). ArXiv preprint arXiv:2402.02330

  19. [27]

    Xu, Z., Yu, C., Fang, F., Wang, Y . & Wu, Y . Language agents with reinforcement learning for strategic play in the werewolf game.Int. Conf. on Mach. Learn. 235, 54889–54910 (2024)

  20. [28]

    Surv.58, 1–41 (2026)

    Mou, X.et al.From individual to society: A survey on social simulation driven by large language model- based agents.ACM Comput. Surv.58, 1–41 (2026)

  21. [29]

    Piao, J.et al.Agentsociety: Large-scale simulation of llm-driven generative agents advances understanding of human behaviors and society.Soc. Sci. Res. Netw. (2025)

  22. [30]

    G.et al.Assessing personality using zero- shot generative ai scoring of brief open-ended text

    Wright, A. G.et al.Assessing personality using zero- shot generative ai scoring of brief open-ended text. Nat. Hum. Behav.1–15 (2026)

  23. [31]

    Serapio-García, G.et al.A psychometric framework for evaluating and shaping personality traits in large language models.Nat. Mach. Intell.1–15 (2025)

  24. [32]

    & Bing, L

    Li, X., Li, Y ., Qiu, L., Joty, S. & Bing, L. Evaluating psychological safety of large language models. In Conference on Empirical Methods in Natural Lan- guage Processing, 1826–1843 (2024)

  25. [33]

    InConference on Empirical Methods in Natural Language Processing, 4125–4143 (2025)

    Wang, S.et al.Exploring the impact of personality traits on llm toxicity and bias. InConference on Empirical Methods in Natural Language Processing, 4125–4143 (2025)

  26. [34]

    M., Vayanos, D

    DeMarzo, P. M., Vayanos, D. & Zwiebel, J. Per- suasion bias, social influence, and unidimensional opinions.The Q. journal economics118, 909–968 (2003)

  27. [35]

    Pragmatics of human communication: A study of interactional patterns, pathologies, and para- doxes.Arch

    Ruesch, J. Pragmatics of human communication: A study of interactional patterns, pathologies, and para- doxes.Arch. Gen. Psychiatry17, 506–507 (1967). 9/26

  28. [36]

    E.The Art of Public Speaking(McGraw- Hill Education, 2015), 12 edn

    Lucas, S. E.The Art of Public Speaking(McGraw- Hill Education, 2015), 12 edn

  29. [37]

    & Keller, K

    Kotler, P. & Keller, K. L.Marketing Management (Pearson Education, 2016), 15 edn

  30. [38]

    Chen, A.et al.The minimax-m2 series: Mini activa- tions unleashing max real-world intelligence.arXiv preprint arXiv:2605.26494(2026). 39.xAI. Grok (2026). Official model page

  31. [40]

    arXiv preprint arXiv:2507.20534(2025)

    Team, K.et al.Kimi k2: Open agentic intelligence. arXiv preprint arXiv:2507.20534(2025)

  32. [41]

    Achiam, J.et al.Gpt-4 technical report.arXiv preprint arXiv:2303.08774(2023)

  33. [42]

    Yang, A.et al.Qwen2.5 technical report.arXiv preprint arXiv:2412.15115(2024)

  34. [43]

    Claude sonnet 4.5 system card

    Anthropic. Claude sonnet 4.5 system card. Tech. Rep., Anthropic (2025). Model system card for Claude Sonnet 4.5

  35. [44]

    Xu, A.et al.Deepseek-v4: Towards highly efficient million-token context intelligence.arXiv preprint arXiv:2606.19348(2026)

  36. [45]

    Team, G.et al.Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805 (2023)

  37. [46]

    Gemini 3.5: Frontier intelligence with action (2026)

    Google. Gemini 3.5: Frontier intelligence with action (2026)

  38. [47]

    The llama 3 herd of models (2024)

    Llama Team. The llama 3 herd of models (2024). ArXiv preprint arXiv:2407.21783

  39. [48]

    stay quiet

    Glm, T.et al.Chatglm: A family of large language models from glm-130b to glm-4 all tools.arXiv preprint arXiv:2406.12793(2024). 49.Seed, B.et al.Seed1. 5-thinking: Advancing superb reasoning models with reinforcement learning.arXiv preprint arXiv:2504.13914(2025). 10/26 Game R...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.