Pith. sign in

REVIEW 101 references

Social Pressure Breaks Majority Voting in LLM Safety Panels

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2608.04415 v1 pith:5TAX4LGE submitted 2026-08-05 cs.CL

classification cs.CL
keywords panelratefalse-alarmmajoritymodelsvotingacrossaverage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but this benefit may disappear when every model sees the same misleading context before voting. We study this risk in a controlled two-round experiment. Each model first judges an item alone, then judges it again after six simulated peers either assert the wrong label or abstain. We combine the final judgments by majority vote. Across six open-weight LLMs and six datasets, we find that the wrong-label peer message raises the average reviewer false-alarm rate from 56.5% under silent peers to 87.5%, and majority voting raises the panel false-alarm rate to 100%. Without an asserted label, the same panel outperforms its average member. The effect is strongly asymmetric: reviewers follow pushes toward "unsafe" far more than pushes toward "safe" (about 75% versus 17%), so the panel's false-alarm rate rises sharply while its harmful-miss rate changes little. The proprietary-model probe shows substantial variation across models. These results identify susceptibility to shared social cues as a failure mode of safety panels and provide a simple pre-deployment diagnostic.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

101 extracted references · 36 canonical work pages

  1. [1]

    Journal of Economic perspectives , volume=

    Cognitive reflection and decision making , author=. Journal of Economic perspectives , volume=. 2005 , publisher=

  2. [2]

    , author=

    Considering the opposite: a corrective strategy for social judgment. , author=. Journal of personality and social psychology , volume=. 1984 , publisher=

  3. [3]

    Scientific american , volume=

    Opinions and social pressure , author=. Scientific american , volume=. 1955 , publisher=

  4. [4]

    , author=

    A study of normative and informational social influences upon individual judgment. , author=. The journal of abnormal and social psychology , volume=. 1955 , publisher=

  5. [5]

    , author=

    Behavioral study of obedience. , author=. The Journal of abnormal and social psychology , volume=. 1963 , publisher=

  6. [6]

    , author=

    The psychology of social impact. , author=. American psychologist , volume=. 1981 , publisher=

  7. [7]

    , author=

    Sources of the continued influence effect: When misinformation in memory affects later inferences. , author=. Journal of experimental psychology: Learning, memory, and cognition , volume=. 1994 , publisher=

  8. [8]

    Psychological science in the public interest , volume=

    Misinformation and its correction: Continued influence and successful debiasing , author=. Psychological science in the public interest , volume=. 2012 , publisher=

Show all 101 references
  1. [9]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Conformity in large language models , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  2. [10]

    arXiv preprint arXiv:2501.13381 , year=

    Do as We Do, Not as You Think: the Conformity of Large Language Models , author=. arXiv preprint arXiv:2501.13381 , year=

  3. [11]

    Herd Behavior: Investigating Peer Influence in

    Cho, Young-Min and Guntuku, Sharath Chandra and Ungar, Lyle , journal=. Herd Behavior: Investigating Peer Influence in

  4. [12]

    When Your

    Mehdizadeh, Aliakbar and Hilbert, Martin , journal=. When Your

  5. [13]

    arXiv preprint arXiv:2601.04790 , year=

    Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework , author=. arXiv preprint arXiv:2601.04790 , year=

  6. [14]

    Conformity Dynamics in

    Han, Chen and Tan, Jin and Yu, Bohan and Zheng, Wenzhen and Tang, Xijin , journal=. Conformity Dynamics in

  7. [15]

    Findings of the Association for Computational Linguistics: ACL 2025 , pages=

    An Empirical Study of Group Conformity in Multi-Agent Systems , author=. Findings of the Association for Computational Linguistics: ACL 2025 , pages=

  8. [16]

    Li, Yuxuan and Guo, Xinwei and Gao, Jiashi and Chen, Guanhua and Zhao, Xiangyu and Zhang, Jiaxin and Liu, Quanying and Wu, Haiyan and Yao, Xin and Wei, Xuetao , booktitle=

  9. [17]

    Justice or Prejudice? Quantifying Biases in

    Jiayi Ye and Yanbo Wang and Yue Huang and Dongping Chen and Qihui Zhang and Nuno Moniz and Tian Gao and Werner Geyer and Chao Huang and Pin-Yu Chen and Nitesh V Chawla and Xiangliang Zhang , booktitle=. Justice or Prejudice? Quantifying Biases in

  10. [18]

    arXiv preprint arXiv:2604.19301 , year=

    Large Language Models Exhibit Normative Conformity , author=. arXiv preprint arXiv:2604.19301 , year=

  11. [19]

    Studies in Social Power , editor=

    The bases of social power , author=. Studies in Social Power , editor=. 1959 , publisher=

  12. [20]

    Sociometry , volume=

    Influence of a consistent minority on the responses of a majority in a color perception task , author=. Sociometry , volume=. 1969 , publisher=

  13. [21]

    PLoS biology , volume=

    Distinct neurocomputational mechanisms support informational and socially normative conformity , author=. PLoS biology , volume=. 2022 , publisher=

  14. [22]

    Organizational behavior and human decision processes , volume=

    Advice taking in decision making: Egocentric discounting and reputation formation , author=. Organizational behavior and human decision processes , volume=. 2000 , publisher=

  15. [23]

    Organizational behavior and human decision processes , volume=

    Trust, confidence, and expertise in a judge-advisor system , author=. Organizational behavior and human decision processes , volume=. 2001 , publisher=

  16. [24]

    Xiong, Miao and Hu, Zhiyuan and Lu, Xinyang and Li, Yifei and Fu, Jie and He, Junxian and Hooi, Bryan , booktitle=. Can

  17. [25]

    International conference on learning representations , volume=

    Large language models cannot self-correct reasoning yet , author=. International conference on learning representations , volume=

  18. [26]

    International Conference on Learning Representations , volume=

    Critic: Large language models can self-correct with tool-interactive critiquing , author=. International Conference on Learning Representations , volume=

  19. [27]

    AutoGen: Enabling Next-Gen

    Qingyun Wu and Gagan Bansal and Jieyu Zhang and Yiran Wu and Beibin Li and Erkang Zhu and Li Jiang and Xiaoyun Zhang and Shaokun Zhang and Jiale Liu and Ahmed Hassan Awadallah and Ryen W White and Doug Burger and Chi Wang , booktitle=. AutoGen: Enabling Next-Gen

  20. [28]

    Forty-first international conference on machine learning , year=

    Improving Factuality and Reasoning in Language Models through Multiagent Debate , author=. Forty-first international conference on machine learning , year=

  21. [29]

    Chan, Chi-Min and Chen, Weize and Su, Yusheng and Yu, Jianxuan and Xue, Wei and Zhang, Shanghang and Fu, Jie and Liu, Zhiyuan , booktitle=

  22. [30]

    Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

    Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration , author=. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techno...

  23. [31]

    Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

    Encouraging divergent thinking in large language models through multi-agent debate , author=. Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

  24. [32]

    Proceedings of the 29th symposium on operating systems principles , pages=

    Efficient memory management for large language model serving with pagedattention , author=. Proceedings of the 29th symposium on operating systems principles , pages=

  25. [33]

    Findings of the Association for Computational Linguistics: ACL 2023 , pages=

    Challenging big-bench tasks and whether chain-of-thought can solve them , author=. Findings of the Association for Computational Linguistics: ACL 2023 , pages=

  26. [34]

    Advances in Neural Information Processing Systems , volume=

    Mmlu-pro: A more robust and challenging multi-task language understanding benchmark , author=. Advances in Neural Information Processing Systems , volume=

  27. [35]

    arXiv preprint arXiv:1803.05457 , year=

    Think you have solved question answering? try arc, the ai2 reasoning challenge , author=. arXiv preprint arXiv:1803.05457 , year=

  28. [36]

    Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

    Truthfulqa: Measuring how models mimic human falsehoods , author=. Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

  29. [37]

    arXiv preprint arXiv:2412.15115 , year=

    Qwen2.5 Technical Report , author=. arXiv preprint arXiv:2412.15115 , year=

  30. [38]

    arXiv preprint arXiv:2310.06825 , year=

    Mistral 7B , author=. arXiv preprint arXiv:2310.06825 , year=

  31. [39]

    arXiv preprint arXiv:2408.00118 , year=

    Gemma 2: Improving Open Language Models at a Practical Size , author=. arXiv preprint arXiv:2408.00118 , year=

  32. [40]

    arXiv preprint arXiv:2407.21783 , year=

    The. arXiv preprint arXiv:2407.21783 , year=

  33. [41]

    arXiv preprint arXiv:2501.00656 , year=

    2 OLMo 2 Furious , author=. arXiv preprint arXiv:2501.00656 , year=

  34. [42]

    International Conference on Learning Representations , volume=

    Darkbench: Benchmarking dark patterns in large language models , author=. International Conference on Learning Representations , volume=

  35. [43]

    arXiv preprint arXiv:2303.13988 , year=

    Machine psychology , author=. arXiv preprint arXiv:2303.13988 , year=

  36. [44]

    Advances in neural information processing systems , volume=

    Self-refine: Iterative refinement with self-feedback , author=. Advances in neural information processing systems , volume=

  37. [45]

    Advances in neural information processing systems , volume=

    Chain-of-thought prompting elicits reasoning in large language models , author=. Advances in neural information processing systems , volume=

  38. [46]

    arXiv preprint arXiv:2605.21318 , year=

    TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization , author=. arXiv preprint arXiv:2605.21318 , year=

  39. [47]

    Advances in neural information processing systems , volume=

    Large language models are zero-shot reasoners , author=. Advances in neural information processing systems , volume=

  40. [48]

    , author=

    Metacognition and cognitive monitoring: A new area of cognitive--developmental inquiry. , author=. American psychologist , volume=. 1979 , publisher=

  41. [49]

    Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining , pages=

    Uncertainty-aware reliable text classification , author=. Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining , pages=

  42. [50]

    Proceedings of the ACM Web Conference 2024 , pages=

    Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries , author=. Proceedings of the ACM Web Conference 2024 , pages=

  43. [51]

    Findings of the Association for Computational Linguistics: EMNLP 2022 , pages=

    Controllable fake document infilling for cyber deception , author=. Findings of the Association for Computational Linguistics: EMNLP 2022 , pages=

  44. [52]

    Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency , pages=

    Understanding the Effects of Explaining Predictive but Unintuitive Features in Human-XAI Interaction , author=. Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency , pages=

  45. [53]

    Proceedings of the 2023 Conference on Human Information Interaction and Retrieval , pages=

    Understanding the cognitive influences of interpretability features on how users scrutinize machine-predicted categories , author=. Proceedings of the 2023 Conference on Human Information Interaction and Retrieval , pages=

  46. [54]

    Proceedings of the 30th ACM International Conference on Information & Knowledge Management , pages=

    A study of explainability features to scrutinize faceted filtering results , author=. Proceedings of the 30th ACM International Conference on Information & Knowledge Management , pages=

  47. [55]

    arXiv:2601.14230 , year=

    MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems , author=. arXiv:2601.14230 , year=

  48. [56]

    Jin, Yiqiao and Zhao, Qinlin and Wang, Yiyang and Chen, Hao and Zhu, Kaijie and Xiao, Yijia and Wang, Jindong , booktitle=

  49. [57]

    2026 , eprint =

    Qu, Jiaming and Fu, Lucheng and Hu, Yibo , title =. 2026 , eprint =

  50. [58]

    2026 , note =

    Hu, Yibo and Qu, Jiaming , title =. 2026 , note =

  51. [59]

    2024 , eprint =

    Xiaochen Zhu and Caiqi Zhang and Tom Stafford and Nigel Collier and Andreas Vlachos , title =. 2024 , eprint =

  52. [60]

    2025 , eprint =

    Huixin Zhong and Yanan Liu and Qi Cao and Shijin Wang and Zijing Ye and Zimu Wang and Shiyao Zhang , title =. 2025 , eprint =

  53. [61]

    2025 , eprint =

    Keyu Wang and Jin Li and Shu Yang and Zhuoran Zhang and Di Wang , title =. 2025 , eprint =

  54. [62]

    2025 , eprint =

    Daniel Vennemeyer and Phan Anh Duong and Tiffany Zhan and Tianyu Jiang , title =. 2025 , eprint =

  55. [63]

    , title =

    Turpin, Miles and Michael, Julian and Perez, Ethan and Bowman, Samuel R. , title =. Advances in Neural Information Processing Systems 36 (NeurIPS) , year =. 2305.04388 , archivePrefix =

  56. [64]

    and Toubia, Olivier , title =

    Brucks, Melanie S. and Toubia, Olivier , title =. PLOS ONE , volume =. 2025 , publisher =

  57. [65]

    Findings of the Association for Computational Linguistics: ACL 2024 , year =

    Madsen, Andreas and Chandar, Sarath and Reddy, Siva , title =. Findings of the Association for Computational Linguistics: ACL 2024 , year =. 2401.07927 , archivePrefix =

  58. [66]

    Strachan, James W. A. and Albergo, Dalila and Borghini, Giulia and Pansardi, Oriana and Scaliti, Eugenio and Gupta, Saurabh and Saxena, Krati and Rufo, Alessandro and Panzeri, Stefano and Manzi, Guido and Graziano, Michael S. A. and Becchio, Cristina , title =. Nature Human Be...

  59. [67]

    and Cheng, Newton and Durmus, Esin and Hatfield-Dodds, Zac and Johnston, Scott R

    Sharma, Mrinank and Tong, Meg and Korbak, Tomasz and Duvenaud, David and Askell, Amanda and Bowman, Samuel R. and Cheng, Newton and Durmus, Esin and Hatfield-Dodds, Zac and Johnston, Scott R. and Kravec, Shauna and Maxwell, Timothy and McCandlish, Sam and Ndousse, Kamal and Ra...

  60. [68]

    Findings of the Association for Computational Linguistics: NAACL 2024 , pages=

    Large language models sensitivity to the order of options in multiple-choice questions , author=. Findings of the Association for Computational Linguistics: NAACL 2024 , pages=

  61. [69]

    2025 , eprint =

    Needham, Joe and Edkins, Giles and Pimpale, Govind and Bartsch, Henning and Hobbhahn, Marius , title =. 2025 , eprint =

  62. [70]

    Findings of the Association for Computational Linguistics: ACL 2023 , month = jul, year =

    Discovering Language Model Behaviors with Model-Written Evaluations , author =. Findings of the Association for Computational Linguistics: ACL 2023 , month = jul, year =. doi:10.18653/v1/2023.findings-acl.847 , pages =

  63. [71]

    arXiv preprint arXiv:2601.03746 , year =

    Schuster, Jakob and Gautam, Vagrant and Markert, Katja , title =. arXiv preprint arXiv:2601.03746 , year =

  64. [72]

    arXiv preprint arXiv:2602.13568 , year =

    Bajaj, Anooshka and Tiganj, Zoran , title =. arXiv preprint arXiv:2602.13568 , year =

  65. [73]

    arXiv preprint arXiv:2601.11563 , year =

    Zhang, Long and Chen, Wei-neng , title =. arXiv preprint arXiv:2601.11563 , year =

  66. [74]

    Journal of Verbal Learning and Verbal Behavior , volume =

    Hasher, Lynn and Goldstein, David and Toppino, Thomas , title =. Journal of Verbal Learning and Verbal Behavior , volume =

  67. [75]

    arXiv preprint arXiv:2508.18321 , year =

    Song, Maojia and Pala, Tej Deep and Zhou, Ruiwen and Jin, Weisheng and Zadeh, Amir and Li, Chuan and Herremans, Dorien and Poria, Soujanya , title =. arXiv preprint arXiv:2508.18321 , year =

  68. [76]

    Mitigating

    Abbasi-Yadkori, Yasin and Kuzborskij, Ilja and Stutz, David and Gy. Mitigating. 2024 , eprint =

  69. [77]

    and Bates, Stephen , title =

    Angelopoulos, Anastasios N. and Bates, Stephen , title =. arXiv preprint arXiv:2107.07511 , year =

  70. [78]

    2025 , eprint =

    Federico Germani and Giovanni Spitale , title =. 2025 , eprint =

  71. [79]

    2023 , eprint =

    Philippe Laban and Lidiya Murakhovs'ka and Caiming Xiong and Chien-Sheng Wu , title =. 2023 , eprint =

  72. [80]

    Science , volume =

    Tversky, Amos and Kahneman, Daniel , title =. Science , volume =

  73. [81]

    Proceedings of the 13th International Conference on Learning Representations (ICLR 2025) , year =

    Jeremy Perez and Grgur Kovac and Corentin Leger and Cedric Colas and Gaia Molinaro and Maxime Derex and Pierre-Yves Oudeyer and Clement Moulin-Frier , title =. Proceedings of the 13th International Conference on Learning Representations (ICLR 2025) , year =. 2407.04503 , archi...

  74. [82]

    Cohen , title =

    Adi Simhi and Fazl Barez and Martin Tutek and Yonatan Belinkov and Shay B. Cohen , title =. 2026 , eprint =

  75. [83]

    Journal of Political Economy , volume =

    Bikhchandani, Sushil and Hirshleifer, David and Welch, Ivo , title =. Journal of Political Economy , volume =

  76. [84]

    , title =

    Kelman, Herbert C. , title =. Journal of Conflict Resolution , volume =

  77. [85]

    2025 , eprint =

    Myra Cheng and Sunny Yu and Cinoo Lee and Pranav Khadpe and Lujain Ibrahim and Dan Jurafsky , title =. 2025 , eprint =

  78. [86]

    2025 , eprint =

    Joshua Liu and Aarav Jain and Soham Takuri and Srihan Vege and Aslihan Akalin and Kevin Zhu and Sean O'Brien and Vasu Sharma , title =. 2025 , eprint =

  79. [87]

    Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI 2026) , year =

    Shomik Jain and Charlotte Park and Matt Viana and Ashia Wilson and Dana Calacci , title =. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI 2026) , year =. 2509.12517 , archivePrefix =

  80. [88]

    2026 , eprint =

    De Marzo, Giordano and Bellina, Alessandro and Castellano, Claudio and Priesemann, Viola and Garcia, David , title =. 2026 , eprint =

  81. [89]

    2024 , eprint =

    Ariel Flint Ashery and Luca Maria Aiello and Andrea Baronchelli , title =. 2024 , eprint =

  82. [90]

    2023 , eprint=

    Ji, Jiaming and Liu, Mickel and Dai, Josef and Pan, Xuehai and Zhang, Chi and Bian, Ce and Chen, Boyuan and Sun, Ruiyang and Wang, Yizhou and Yang, Yaodong , booktitle=. 2023 , eprint=

  83. [91]

    Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

    R. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=. 2024 , eprint=

  84. [92]

    2024 , eprint=

    Zhang, Zhexin and Lei, Leqi and Wu, Lindong and Sun, Rui and Huang, Yongkang and Long, Chong and Liu, Xiao and Lei, Xuanyu and Tang, Jie and Huang, Minlie , booktitle=. 2024 , eprint=

  85. [93]

    Proceedings of the IEEE/ACM International Conference on Software Engineering (ICSE) , year=

    Vulnerability Detection with Code Language Models: How Far Are We? , author=. Proceedings of the IEEE/ACM International Conference on Software Engineering (ICSE) , year=. 2403.18624 , archivePrefix=

  86. [94]

    Aligning

    Hendrycks, Dan and Burns, Collin and Basart, Steven and Critch, Andrew and Li, Jerry and Song, Dawn and Steinhardt, Jacob , booktitle=. Aligning. 2021 , eprint=

  87. [95]

    2024 , eprint=

    Han, Seungju and Rao, Kavel and Ettinger, Allyson and Jiang, Liwei and Lin, Bill Yuchen and Lambert, Nathan and Choi, Yejin and Dziri, Nouha , booktitle=. 2024 , eprint=

  88. [96]

    Ghosh, Shaona and Varshney, Prasoon and Galinkin, Erick and Parisien, Christopher , journal=

  89. [97]

    Inan, Hakan and Upasani, Kartikeya and Chi, Jianfeng and Rungta, Rashi and Iyer, Krithika and Mao, Yuning and Tontchev, Michael and Hu, Qing and Fuller, Brian and Testuggine, Davide and Khabsa, Madian , journal=

  90. [98]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    A Holistic Approach to Undesired Content Detection in the Real World , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=. doi:10.1609/aaai.v37i12.26752 , year=

  91. [99]

    Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

    Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive Contexts , author=. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=. 2021 , eprint=

  92. [100]

    , title =

    Ladha, Krishna K. , title =. American Journal of Political Science , volume =

  93. [101]

    2024 , eprint =

    Verga, Pat and Hofstatter, Sebastian and Althammer, Sophia and Su, Yixuan and Piktus, Aleksandra and Arkhangorodsky, Arkady and Xu, Minjie and White, Naomi and Lewis, Patrick , title =. 2024 , eprint =

Pith tools