REVIEW 3 major objections 6 minor 53 references
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read LLM agents barely reproduce human pronoun patterns in simulated team interactions, even when they know the patterns.
desk verdict A useful cautionary study on LLM social simulation, but the abstract overclaims and the round-robin protocol may suppress the signal worth studying. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is the leader/non-leader pronoun-frequency differential, Δ = f_nonleaders,avg − f_leaders,avg, computed for first-person singular and first-person plural pronouns across 41 simulated four-person groups; paired with four persona prompts from prior work and a reflection/planning agent adapted from Park et al. (2023), it converts a well-known human interaction norm into a quantitative, falsifiable benchmark for agent behavior.
What would settle it
A replication that lets four LLM agents hold a free-form, 30-minute discussion of the same ranking task and finds a statistically significant positive Δ for first-person singular pronouns (non-leaders above leaders) and negative Δ for first-person plural across most models and prompts would contradict the paper's central claim; so would a controlled test showing that the three-round round-robin format itself suppresses human-like pronoun patterns when humans speak under the same constraint.
Extended reading notes
Core claim
The paper's central claim is that LLM agents barely demonstrate human-like pronoun patterns, even if the LLM agent may show some understanding of those patterns. Replicating the four-person, task-oriented ranking discussion of Kacewicz et al. (2014) with 41 simulated groups, the authors measure the difference in pronoun frequency between non-leaders and leaders, Δ = f_nonleaders,avg − f_leaders,avg, for first-person singular and first-person plural pronouns. Contrary to human results, most non-leader LLMs do not use first-person singular pronouns more often, and leader LLMs do not use first-person plural pronouns more often; in many cases the trend is reversed, and almost none of the tested models, prompts, or the reflection/planning specialized agent yield statistically significant human-like differences. The paper also establishes a knowledge–demonstration gap: several LLMs, when asked directly and with permuted answer orders, can identify the correct human pattern but fail to exhibit it in their interaction process.
Load-bearing premise
The simulation is a fair stand-in for the human study: three rounds of round-robin turns among four agents produce pronoun frequencies comparable to the original 30-minute human group discussions.
Editorial extensions
If this is right
- Conclusions drawn from LLM social simulations about group interaction processes may contradict established human findings, so practitioners should not rely on these simulations for decisions about real social dynamics.
- Adding reflection and planning components does not close the gap; for GPT-4o it actually made pronoun patterns less human-like than the simple prompt-based agent.
- Model family and prompt choice, not model size or capability, dominate pronoun patterns: Llama and Qwen families show internally consistent but non-human trends, and larger models are not more human-like.
- The knowledge–demonstration gap means that an LLM's ability to state a social norm is not evidence that its multi-agent interactions embody that norm.
- Even explicit instructions to use certain pronouns more often (e.g., "Please use first-person plural forms more often" for leaders) did not elicit the human pattern.
Reading between the lines
- If the gap generalizes beyond pronouns, other unconscious conversational signals such as turn-taking, interruption, politeness, and hedging may also fail to emerge in LLM agent interactions, so interaction-level validation should become standard before simulation results are used.
- The three-round, round-robin design may compress the dynamics that let human leaders and subordinates negotiate roles; longer or free-form conversations, or letting agents choose turn order, might allow status roles to crystallize and should be tested before concluding the failure is intrinsic.
- The knowledge–demonstration gap suggests a testable intervention: prompting agents to explicitly monitor their own pronoun use or reflect on their role after each turn might close the gap, though the paper's planning/reflection results hint the opposite, pointing to a dissociation between declarative social knowledge and procedural language generation.
- Another testable extension is to measure whether the same failure appears consistently across random seeds for a fixed model and prompt, which would separate genuine role-behavior deficits from generation variance.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper asks whether multi-agent LLM simulations reproduce a well-known human result on pronoun usage in hierarchical groups (Kacewicz et al., 2014). The authors simulate 41 four-person groups with one randomly assigned leader and three subordinates, using round-robin, three-turn task-oriented discussions. They test GPT, Llama, Mistral, and Qwen models with four persona prompts from the literature, plus a reflection/planning variant of GPT-4o, and compare the non-leader minus leader frequency differences for first-person singular and first-person plural pronouns against the human benchmark. They find that most model-prompt combinations do not show a statistically significant difference in the human direction; GPT-4o with Prompts 2-4 and GPT-4 with several prompts are exceptions. A separate knowledge probe shows that many of the same LLMs can answer which role tends to use a given pronoun more often, despite failing to reproduce the pattern in interaction. The paper concludes that LLM agents barely resemble human interaction processes and urges caution in LLM-based social simulation.
Significance. If the main result holds, the paper is a useful cautionary contribution to LLM-based social simulation: it is one of the few studies that evaluates an interaction process (pronoun frequencies) rather than only the outcome, it uses an external psychology benchmark from the outset, and it reports full per-condition tables in the appendix, making the evaluation transparent. The knowledge-versus-demonstration gap is a genuine and thought-provoking discrepancy. However, the central negative claim is currently stronger than the evidence, because the simulation protocol may suppress the conversational contingencies that produce the human asymmetry and because null findings are interpreted without a statistical power analysis.
major comments (3)
- [§3.4, Algorithm 1] The round-robin protocol in Algorithm 1, with three turns per agent, produces long self-contained monologues rather than the contingent multi-party talk in Kacewicz et al. (2014); Table 9 and Table 10 show that even the non-leader turns are polite collaborative speeches. The argument in §3.4 that 'the frequency does not rely on the number of words generated from each agent' addresses only the unbiasedness of a proportion, not whether the conversational affordances that generate the human asymmetry are present. Since the human effect is theorized to arise in reactive, short back-and-forth exchanges, the negative results on most models may be an artifact of the protocol. I ask the authors to provide evidence that the pronoun asymmetry survives in human data under the same long-turn, fixed-order protocol, or to demonstrate it with a more natural interaction regime before claiming that LLM agents cannot replicate human interactions.
- [§4.2, Eq. (1), Figures 2-6] The evaluation treats any non-significant difference as a failure (gray bars), but the human effect sizes are small (approximately 1.3 percentage points for first-person singular and approximately -0.5 for first-person plural, Table 4a). With 41 groups and the per-condition variation shown in Appendix B, many of the gray bars may simply reflect low statistical power, and the paper does not report effect sizes, confidence intervals, or a power analysis. The central conclusion that LLM agents barely demonstrate human-like pronoun patterns is therefore stronger than the statistics warrant; a power analysis or an equivalence-testing framework is needed to support the null results.
- [Abstract, §5.1, §5.3, Conclusion] The abstract and conclusion state that prompt-based and specialized agents 'fail to demonstrate human-like pronoun usage patterns,' but Figure 2a shows GPT-4o with Prompts 2-4 reproducing the human first-person singular pattern, and Figure 5b shows GPT-4 with Prompts 1, 2, and 4 reproducing the human first-person plural pattern. The evidence supports 'most settings fail' but not a categorical failure; the prose should be calibrated to the observed success rate of approximately 8 human-like significant conditions out of 112 model-prompt-pronoun combinations in Table 2.
minor comments (6)
- [§4.2] Please specify whether the two-sample t-test is paired or unpaired; since leaders and non-leaders come from the same 41 groups, a paired test is the natural choice but is not stated.
- [§4.1] The 'May 13th 2025 version' of GPT-4o is inconsistent with the January 2025 submission date and should be corrected.
- [§5.4, Table 2] A sentence explaining how the 'Dem.' counts are derived (number of persona prompts, out of four, with a statistically significant human-like difference) would aid the reader.
- [Appendix B.2, Table 7] The Llama 3.1 70B rows list 'Prompt 2' twice, with the second 'Prompt 3' row containing Prompt 2 values; this appears to be a copy-paste error.
- [§5.4] Calling the forced-choice questionnaire 'know' overstates what is probed; the prompt asks for the aggregate frequency fact, not the context-sensitive sociolinguistic knowledge required for spontaneous use.
- [Figures 2-6] The red check marks described in Section 5 are faint in the reproduced figures; please make them more visible or use a different marker.
Circularity Check
No circularity: the benchmark is an external psychology result, no parameters are fitted to reproduce it, and the knowledge probe is an independent test.
full rationale
The paper's central claim is an empirical negative result: LLM agents, under the prompts and agent designs tested, do not reproduce the human leader/non-leader pronoun asymmetries reported by Kacewicz et al. (2014). The benchmark is external and prior, not derived from the paper's own outputs or fitted to the LLM data. The evaluation statistic (Equation 1, the difference in averaged pronoun frequencies) is defined from the LLM transcripts and compared against the human benchmark, but no model parameter or prompt is tuned to force agreement with the human result; in fact, the paper reports failures across most configurations. The knowledge probe in Section 5.4 is a separate query task, so the 'know vs. demonstrate' distinction is not constructed from the interaction data. The paper's self-citations (e.g., Borah and Mihalcea 2024 for a persona prompt, Wu et al. 2023 for a similar knowledge-behavior disparity, Piatti et al. 2024 for prior simulation work) are contextual and none is load-bearing for the main conclusion. The defense in Section 3.4 that pronoun frequency ratios should not depend on number of words is a claim about estimation bias, not an equation that defines the prediction in terms of the input; even if that claim is debatable as a matter of experimental validity or statistical power, it is not circularity. No fitted input is renamed as a prediction, no uniqueness theorem is imported from the authors, and no ansatz is smuggled in via citation. The paper is self-contained against an external benchmark, so the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (2)
- number_of_interaction_rounds =
3
- sampling_temperature =
0.7
assumptions (5)
- domain assumption Kacewicz et al. (2014)'s human pronoun results are a valid gold standard: non-leaders use more first-person singular pronouns and leaders use more first-person plural pronouns in hierarchical group discussions.
- ad hoc to paper Pronoun usage frequency is independent of the number of words generated per agent, so three rounds of interaction are equivalent to the original 30-minute human sessions for measuring this pattern.
- ad hoc to paper Round-robin turn-taking does not alter the leader versus non-leader pronoun pattern compared to natural human turn-taking.
- domain assumption The four persona prompts P1-P4 are representative of how LLM social simulations are configured in practice.
- standard math A two-sample t-test on 41 leader and 41 non-leader frequency values is a valid significance test for the difference.
Cite this review
Pith. "Pith review of Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions." pith.science (2026). https://pith.science/paper/5PYGJ6FZ
@misc{pith2026250115283,
author = {Pith},
title = {Pith review of: Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions},
year = {2026},
howpublished = {\url{https://pith.science/paper/5PYGJ6FZ}},
note = {Machine review of arXiv:2501.15283}
}
read the original abstract
As Large Language Models (LLMs) advance in their capabilities, researchers have increasingly employed them for social simulation. In this paper, we investigate whether interactions among LLM agents resemble those of humans. Specifically, we focus on the pronoun usage difference between leaders and non-leaders, examining whether the simulation would lead to human-like pronoun usage patterns during the LLMs' interactions. Our evaluation reveals the significant discrepancies between LLM-based simulations and human pronoun usage, with prompt-based or specialized agents failing to demonstrate human-like pronoun usage patterns. In addition, we reveal that even if LLMs understand the human pronoun usage patterns, they fail to demonstrate them in the actual interaction process. Our study highlights the limitations of social simulations based on LLM agents, urging caution in using such social simulation in practitioners' decision-making process.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[2]
Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai. 2023. Using large language models to simulate multiple humans and replicate human subject studies. In International Conference on Machine Learning, pages 337--371. PMLR
2023
-
[3]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609
arXiv 2023
-
[4]
John A Bargh. 1989. Conditional automaticity: Varieties of automatic influence in social perception and cognition. Unintended thought, pages 3--51
work page 1989
-
[5]
John A Bargh, Kay L Schwader, Sarah E Hailey, Rebecca L Dyer, and Erica J Boothby. 2012. Automaticity in social-cognitive processes. Trends in cognitive sciences, 16(12):593--605
work page 2012
-
[6]
Angana Borah and Rada Mihalcea. 2024. http://arxiv.org/abs/2410.02584 Towards implicit bias detection and mitigation in multi-agent llm interactions
arXiv 2024
-
[7]
Marilynn B Brewer and Wendi Gardner. 1996. Who is this" we"? levels of collective identity and self representations. Journal of personality and social psychology, 71(1):83
work page 1996
-
[8]
Ethan R Burris, Matthew S Rodgers, Elizabeth A Mannix, Michael G Hendron, and James B Oldroyd. 2009. Playing favorites: The influence of leaders' inner circle on group processes and performance. Personality and Social Psychology Bulletin, 35(9):1244--1257
work page 2009
Show all 53 references
-
[9]
R Sherlock Campbell and James W Pennebaker. 2003. The secret life of pronouns: Flexibility in writing style and physical health. Psychological science, 14(1):60--65
2003
-
[10]
Shuaichen Chang and Eric Fosler-Lussier. 2023. How to prompt llms for text-to-sql: A study in zero-shot, single-domain, and cross-domain settings. arXiv preprint arXiv:2305.11853
2023 arXiv
-
[11]
Deborah Davis and Timothy C Brock. 1975. Use of first person pronouns as a function of increased objective self-awareness and performance feedback. Journal of Experimental Social Psychology, 11(4):381--388
1975
-
[12]
Naihao Deng, Zhenjie Sun, Ruiqi He, Aman Sikka, Yulong Chen, Lin Ma, Yue Zhang, and Rada Mihalcea. 2024. https://doi.org/10.18653/v1/2024.findings-acl.23 Tables as texts or images: Evaluating the table reasoning ability of LLM s and MLLM s . In Findings of the Association for ...
2024 doi
-
[13]
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.88 Toxicity in chatgpt: Analyzing persona-assigned language models . In Findings of the Association for Computational Linguistics:...
2023 doi
-
[14]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[15]
Shelley Duval and Robert A Wicklund. 1972. A theory of objective self awareness
1972
-
[16]
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323
2022 arXiv
-
[17]
Chen Gao, Xiaochong Lan, Zhihong Lu, Jinzhu Mao, Jinghua Piao, Huandong Wang, Depeng Jin, and Yong Li. 2023. S ^3 : Social-network simulation system with large language model-empowered agents. arXiv preprint arXiv:2307.14984
2023 arXiv
-
[18]
John J Gumperz. 1982. Language and social identity. Cambridge University Press
1982
-
[19]
Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. 2023. Bias runs deep: Implicit reasoning biases in persona-assigned llms. arXiv preprint arXiv:2311.04892
2023 arXiv
-
[20]
Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji, Yingqiang Ge, Libby Hemphill, and Yongfeng Zhang. 2023. War and peace (waragent): Large language model-based multi-agent simulation of world wars. arXiv preprint arXiv:2311.17227
2023 arXiv
-
[21]
Molly E Ireland and Matthias R Mehl. 2014. Natural language use as a marker. The Oxford handbook of language and social psychology, pages 201--237
2014
-
[22]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 a . Mistral 7b. arXiv preprint arXiv:2310.06825
2023 arXiv
-
[23]
Hang Jiang, Xiajie Zhang, Xubo Cao, Jad Kabbara, and Deb Roy. 2023 b . Personallm: Investigating the ability of gpt-3.5 to express personality traits and gender differences. arXiv preprint arXiv:2305.02547
2023 arXiv
-
[24]
Mingyu Jin, Beichen Wang, Zhaoqian Xue, Suiyuan Zhu, Wenyue Hua, Hua Tang, Kai Mei, Mengnan Du, and Yongfeng Zhang. 2024. What if llms have different world views: Simulating alien civilizations with llm-based agents. arXiv preprint arXiv:2402.13184
2024 arXiv
-
[25]
Ewa Kacewicz, James W Pennebaker, Matthew Davis, Moongee Jeon, and Arthur C Graesser. 2014. Pronoun use reflects standings in social hierarchies. Journal of Language and Social Psychology, 33(2):125--143
2014
-
[26]
Zhao Kaiya, Michelangelo Naim, Jovana Kondic, Manuel Cortes, Jiaxin Ge, Shuying Luo, Guangyu Robert Yang, and Andrew Ahn. 2023. Lyfe agents: Generative agents for low-cost real-time social interactions. arXiv preprint arXiv:2310.02172
2023 arXiv
-
[27]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Princip...
2023
-
[28]
Kenneth Li, Tianle Liu, Naomi Bashkansky, David Bau, Fernanda Vi \'e gas, Hanspeter Pfister, and Martin Wattenberg. 2024 a . Measuring and controlling persona drift in language model dialogs. arXiv preprint arXiv:2402.10962
2024 arXiv
-
[29]
Nian Li, Chen Gao, Mingyu Li, Yong Li, and Qingmin Liao. 2024 b . https://doi.org/10.18653/v1/2024.acl-long.829 E con A gent: Large language model-empowered agents for simulating macroeconomic activities . In Proceedings of the 62nd Annual Meeting of the Association for Comput...
2024 doi
-
[30]
Yuan Li, Yixuan Zhang, and Lichao Sun. 2023. Metaagents: Simulating interactions of human behaviors for llm-based task-oriented coordination via collaborative generative agents. arXiv preprint arXiv:2310.06500
2023 arXiv
-
[31]
Matthias R Mehl, Samuel D Gosling, and James W Pennebaker. 2006. Personality in its natural habitat: manifestations and implicit folk theories of personality in daily life. Journal of personality and social psychology, 90(5):862
2006
-
[32]
Joon Sung Park, Joseph O'Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, pages 1--22
2023
-
[33]
Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2022. Social simulacra: Creating populated prototypes for social computing systems. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Techno...
2022
-
[34]
Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. 2024. Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109
2024 arXiv
-
[35]
James W Pennebaker. 2011. The secret life of pronouns. New Scientist, 211(2828):42--45
2011
-
[36]
Alastair Pennycook. 1994. The politics of pronouns
1994
-
[37]
Giorgio Piatti, Zhijing Jin, Max Kleiman-Weiner, Bernhard Sch \"o lkopf, Mrinmaya Sachan, and Rada Mihalcea. 2024. Cooperate or collapse: Emergence of sustainability behaviors in a society of llm agents. arXiv preprint arXiv:2404.16698
2024 arXiv
-
[38]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9
2019
-
[39]
Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. 2024. In-context impersonation reveals large language models' strengths and biases. Advances in Neural Information Processing Systems, 36
2024
-
[40]
Jan E Stets and Chris F Biga. 2003. Bringing identity theory into environmental sociology. Sociological theory, 21(4):398--423
2003
-
[41]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971
2023 arXiv
-
[42]
Mike Treanor, Ben Samuel, and Mark J Nelson. 2024. Prototyping slice of life: Social physics with symbolically grounded llm-based generative dialogue. In Proceedings of the 19th International Conference on the Foundations of Digital Games, pages 1--4
2024
-
[43]
Noah Wang, Z.y. Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, and Junran Peng. 2024 a . https://doi.org/10.18653/v1/2024.findings-acl.878 R ole ...
2024 doi
-
[44]
Shuai Wang, Harrisen Scells, Bevan Koopman, and Guido Zuccon. 2023. Can chatgpt write a good boolean query for systematic review literature search? In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1426--1436
2023
-
[45]
Yudong Wang, Damai Dai, and Zhifang Sui. 2024 b . Exploring activation patterns of parameters in language models. arXiv preprint arXiv:2405.17799
2024 arXiv
-
[46]
Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. 2023. A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382
2023 arXiv
-
[47]
Yufan Wu, Yinghui He, Yilin Jia, Rada Mihalcea, Yulong Chen, and Naihao Deng. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.717 Hi- T o M : A benchmark for evaluating higher-order theory of mind reasoning in large language models . In Findings of the Association for Co...
2023 doi
-
[48]
Benfeng Xu, An Yang, Junyang Lin, Quan Wang, Chang Zhou, Yongdong Zhang, and Zhendong Mao. 2023. Expertprompting: Instructing large language models to be distinguished experts. arXiv preprint arXiv:2305.14688
2023 arXiv
-
[49]
Zhengyan Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2022. https://doi.org/10.18653/v1/2022.findings-acl.71 M o E fication: Transformer feed-forward layers are mixtures of experts . In Findings of the Association for Computational Linguistics: ACL 2022,...
2022 doi
-
[50]
Zheyuan Zhang, Daniel Zhang-Li, Jifan Yu, Linlu Gong, Jinchang Zhou, Zhiyuan Liu, Lei Hou, and Juanzi Li. 2024. Simulating classroom education with llm-empowered agents. arXiv preprint arXiv:2406.19226
2024 arXiv
-
[51]
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2023. Large language models are not robust multiple choice selectors. In The Twelfth International Conference on Learning Representations
2023
-
[52]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...
-
[53]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.