REVIEW 3 major objections 4 minor 51 references
Measuring Human Leadership Skills with Artificially Intelligent Agents
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A new experiment claims that a leader's skill with AI agents predicts their causal impact on human teams, with a disattenuated correlation of 0.81.
desk verdict A serious, pre-registered measurement study with a real and interesting correlation, but the abstract overclaims that AI agents are effective proxies for human participants without validating follower behavior. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the AI leadership test built on a modified Hidden Profile paradigm: a leader and three followers each hold private clues, the leader can talk to anyone but followers can only talk to the leader, and the leader must synthesize dispersed information into probabilistic answers. The ground-truth measure is the leader's average causal contribution estimated by randomly assigning each leader to multiple human groups and then using a multilevel model to isolate the leader-effect standard deviation. The paper's key comparison is between leader-effect estimates from the AI version and the human version, with parallel puzzle forms created so that the AI agents would not have seen solutions in training data.
What would settle it
Re-run the same pre-registered design with a different large language model or with follower prompts that reduce cooperativeness or information sharing; if the disattenuated correlation with human ground-truth causal contributions drops well below 0.81, or if the leader-effect sizes diverge, the proxy is specific to the original agent configuration rather than to general leadership skill.
Extended reading notes
Core claim
The paper's central discovery is that individual scores on an AI leadership test, where a human directs three LLM followers through collaborative puzzles, strongly predict the same leader's total causal contribution to teams of human followers. Using a hidden-profile task with a distinct leader role and probabilistic answers, the authors randomly assigned each leader to six human groups and six AI-agent groups. The disattenuated correlation between AI-test scores and human ground-truth scores is rho = 0.81 (95% CI [0.72, 0.88]); after conditioning on leader hard skills such as fluid IQ, typing speed, and task-specific ability, the correlation remains 0.69 (95% CI [0.57, 0.81]). The same leader characteristics predict success in both settings, and the same communication behaviors (asking questions, conversational turn-taking, using 'we' language) are associated with performance. The authors interpret this as evidence that LLM agents can be effective proxies for human participants in measuring leadership and teamwork skills.
Load-bearing premise
The result stands or falls on whether LLM followers, prompted with instructions analogous to those given to human followers, respond to leader behavior through the same causal channels (information pooling, question asking, turn-taking) that make leaders effective with human teams.
Editorial extensions
If this is right
- Leadership assessment becomes scalable: the AI version cost about $23 per leader and ran autonomously, versus $114 and live coordination for the human version, so repeated-random-assignment measurement could be used much more widely.
- The AI test predicts not just raw leadership success but also leadership-specific soft skills, since the correlation survives conditioning on hard skills (disattenuated rho = 0.69).
- Substantive findings about leadership, such as overconfidence predicting willingness to lead and accurate self-evaluators contributing more, replicate in the AI test, suggesting silicon samples can test leadership hypotheses.
- Because demographics do not predict leadership skill, a cheap standardized AI-based test could in principle identify high-potential leaders overlooked by current selection processes.
Reading between the lines
- If the proxy relationship holds, the same repeated-random-assignment design could measure causal leader contributions longitudinally, enabling pre/post evaluation of leadership training programs that currently lack robust outcome measures.
- The paper's own evidence that positive affect matters less in the AI test suggests that AI-follower assessments may underweight socioemotional leadership; a test battery that deliberately varies follower affect or cooperativeness could recover that dimension.
- The raw (non-disattenuated) correlation is 0.67, so in practical selection settings with short, noisy assessments the predictive validity will be lower than the headline 0.81; operational use would need multiple sessions or longer tests.
- The hidden-profile task is one specific team process; whether the human-AI correlation generalizes to other team activities, such as creative brainstorming or crisis decision-making, remains an open testable question.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a pre-registered lab experiment (n=249 leaders) in which each leader directed teams of three human followers and, separately, teams of three GPT-4o AI followers on a modified Hidden-Profile task. Leaders were randomly assigned to multiple human teams to estimate their causal contribution to group performance, and this ground-truth measure was compared with the same leaders' performance on an AI version of the task. The authors report a raw leader-level correlation of 0.67 between AI and human test performance, rising to a disattenuated correlation of 0.81, and similar patterns of predictors (skill measures, demographics) and communication behaviors across the two tests. They conclude that AI agents can serve as effective proxies for human participants in social experiments, enabling scalable measurement of leadership skills.
Significance. If the central correlation is robust and generalizable, the paper offers a low-cost, performance-based method for measuring leadership skill, with implications for personnel selection, leadership training, and experimental social science. The design has notable strengths: it is pre-registered, the order of AI and human tests is counterbalanced, leaders are repeatedly randomly assigned to human teams, hard-skill confounders (typing speed, individual hidden-profile skill, fluid IQ) are measured and conditioned on, process measures (question asking, turn-taking, pronoun use) are extracted from chat logs, and the data, code, and materials are promised publicly. The reported correlations are large and the predictor profiles are similar across tests. However, the headline disattenuated correlation is not fully transparent, the equivalence of the two puzzle forms is asserted rather than demonstrated, and the 'AI agents as proxies' claim goes beyond the leader-level evidence provided.
major comments (3)
- [Section 3, 'Individual scores on the AI test correlate very highly with ground-truth scores'] The headline estimate ρ̂=0.81 is a disattenuated correlation, but the manuscript never specifies the reliability coefficients or the attenuation-correction formula used. Without this information the reader cannot judge whether 0.81 is an over-correction, which is load-bearing for the paper's main claim. Please report the reliability estimates (e.g., split-half or intraclass correlations) for Y_i^AI and Y_i^Human, the exact correction formula, and a sensitivity analysis using alternative reliability values. Also, the confidence interval is printed as '[0.72,88]' and should read '[0.72,0.88]'.
- [Section 2, 'Group task'] The claim that the two parallel puzzle forms have 'equivalent structure and difficulty across items' is asserted without supporting evidence. If the human and AI forms differ in difficulty or in the distribution of clue types, the leader-level correlation between the two tests could be inflated or attenuated. Because every leader took both forms, a form-order confound would directly affect the cross-test correlation. Please provide evidence of form equivalence (e.g., pilot calibration data or pretest statistics) and clarify whether the forms were randomly assigned or counterbalanced with respect to the human/AI test order and the two-order condition.
- [Abstract and Discussion] The conclusion that 'AI agents can be effective proxies for human participants in social experiments' goes beyond the evidence presented. The study validates the AI test at the level of leader outcomes (high correlation with human-team leadership, similar predictor profiles), but it does not demonstrate that AI followers behave like human followers: the chat logs are not analyzed to compare information disclosure, misinterpretation, compliance, or emotional reactions. The authors' own results show that positive affect does not predict AI-team performance and that emotional perceptiveness is a weaker predictor of AI-test success, which indicates that the AI followers are not fully human-like. To support the proxy claim, either the conclusions should be scoped to 'AI-based assessment captures leader skills that transfer to human teams,' or the authors should add follower-level behavioral comparisons from the chat data. Without such evidence, the 0.81 correlation could partly reflect common cognitive skill (e.g., verbal ability or executive function) rather than social leadership, and conditioning on measured hard skills (Eqs. 1-2) does not eliminate this concern because unmeasured abilities could still load on both tests.
minor comments (4)
- [Methods, Eqs. (1)-(2)] The subscripts in equation (1) are inconsistent: the sum is written as Σ_g I_{ig} X_i, but the left-hand side is indexed by g; this should likely be Σ_i I_{ig} X_i, with the residual indexed by g (ε_g) rather than ε_{ig}. Please clarify the notation so the definitions of α_i in equation (2) are unambiguous.
- [Discussion] The sentence 'mirroring findings from the field (8)' cites reference 8 (Argyle et al., on simulating human samples), but the intended reference is presumably to the empirical literature on gender, ethnicity, and leadership (e.g., references 30–32). Please correct this citation.
- [Figure 5 caption] The caption describes the left panel as showing 'Leadership Skill,' whereas the text says the left panel conditions on task-specific skill. Please make the caption consistent with the text (e.g., 'leadership skill after conditioning on hard skills').
- [Abstract and main text] The term 'disattenuated correlation' is used without a parenthetical explanation. Consider adding a brief note defining the correction for measurement error, as many readers will not be familiar with psychometric terminology.
Circularity Check
No significant circularity: the central claim is an empirical correlation between an AI-based test and a human ground-truth measure, with no fitted parameters renamed as predictions and no load-bearing self-citation chain.
full rationale
The paper's central claim is empirical, not derivational. Human leaders are randomly assigned to teams of human followers to obtain a ground-truth causal contribution, and separately lead teams of GPT-4o agents on newly created puzzles; the paper then reports the correlation between the two measures. No equation in the paper defines the AI test score in terms of the human test score, and no parameter is fitted to the human outcome and then presented as a prediction of it. The disattenuated correlation (ρ̂=0.81) is an estimate of association, not a fitted constant that reproduces the input by construction. The AI agents are prompted with instructions analogous to those given to human participants, but the puzzles are newly created so the agents could not have memorized the solutions, and the comparison against human teams provides external benchmark data. The paper acknowledges limitations, including the diminished role of emotion and the fact that AI agents do not replicate the full variety of human behavior; these are validity concerns, not circularity. Some self-citations appear (the multilevel model from Weidmann and Deming 2021 and the PAGE emotion perception measure from Weidmann and Xu 2024), but they are used as measurement instruments or statistical tools, not as the source of the central claim, and the central correlation is independently computed from the experiment's own data. Therefore no circular step can be exhibited, and the appropriate score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption Random assignment of leaders to groups identifies causal leader effects.
- domain assumption GPT-4o agents behave analogously to human followers in the hidden profile task.
- domain assumption The human and AI versions of the hidden profile puzzles are parallel forms of equivalent difficulty.
- domain assumption Residual leader effects after conditioning on measured hard skills capture leadership soft skills.
- standard math Normality and homoskedasticity of random effects in the multilevel model.
Cite this review
Pith. "Pith review of Measuring Human Leadership Skills with Artificially Intelligent Agents." pith.science (2026). https://pith.science/paper/7JPLJVPU
@misc{pith2026250802966,
author = {Pith},
title = {Pith review of: Measuring Human Leadership Skills with Artificially Intelligent Agents},
year = {2026},
howpublished = {\url{https://pith.science/paper/7JPLJVPU}},
note = {Machine review of arXiv:2508.02966}
}
read the original abstract
We show that the ability to lead groups of humans is predicted by leadership skill with Artificially Intelligent agents. In a large pre-registered lab experiment, human leaders worked with AI agents to solve problems. Their performance on this 'AI leadership test' was strongly correlated with their causal impact on human teams, which we estimate by repeatedly randomly assigning leaders to groups of human followers and measuring team performance. Successful leaders of both humans and AI agents ask more questions and engage in more conversational turn-taking; they score higher on measures of social intelligence, fluid intelligence, and decision-making skill, but do not differ in gender, age, ethnicity or education. Our findings indicate that AI agents can be effective proxies for human participants in social experiments, which greatly simplifies the measurement of leadership and teamwork skills.
Reference graph
Works this paper leans on
-
[1]
Variation in leader performance has a large and statistically significant impact on group performance, for both the AI test and the human ground truth
-
[2]
The same skills and demographic factors predict success on both tests
-
[3]
This is true for raw scores and leadership measures that control for hard skills
Individual scores on the AI test correlate very highly with ground-truth scores. This is true for raw scores and leadership measures that control for hard skills. We then move to exploratory analysis and show that:
-
[4]
The same behavioral patterns predict success in both tests
-
[5]
The AI Leadership Test reproduces substantive findings about leadership emergence and performance
-
[6]
Variation in leader performance matters for group performance in both tests We find strong evidence that ‘leader effects’ – i.e. the causal impact a typical leader has on team performance – are large for both the AI leadership test and the human test (see Method section for formal definition of leader effects). More than half the variation in group perfor...
-
[7]
The same skills/characteristics predict success on the AI test and the ground-truth If the AI Leadership Test does a good job of replicating the ground-truth test, we would expect that similar leader characteristics and skill profiles would predict success on both tests. This is what we find. Figure 5 summarizes these results: the x-axis shows the correla...
-
[8]
Individual scores on the AI test correlate very highly with ground-truth scores We examine each leader’s average raw score on the AI leadership test (𝑌𝑖 𝐴𝐼) and their average score across groups on the human test (𝑌𝑖 𝐻𝑢𝑚𝑎𝑛). At the level of individual leaders, the disattenuated correlation is 𝜌̂ =0.81 (n=249), with a 95% confidence interval of [0.72,88].7...
Show all 51 references
-
[9]
We start by analyzing the communicative strategies that are associated with success on the AI and Human leadership tests
The same behavioral patterns are associated with success in AI and Human tests Next, we turn to exploratory analyses. We start by analyzing the communicative strategies that are associated with success on the AI and Human leadership tests. We explore five process metrics, meas...
-
[10]
We now explore whether the AI test may also benefit social science by testing substantive hypotheses about leadership and teamwork dynamics
The AI Leadership Test reproduces substantive findings about leadership Thus far, we’ve shown that individual scores on the AI Leadership Test are remarkably similar to the ground-truth scores. We now explore whether the AI test may also benefit social science by testing subst...
-
[11]
The task, described in detail in the next section, measures each leader’s ability to gather information and make decisions
Participant recruitment and flow The core of our experiment is a collaborative problem-solving task based on the Hidden Profile paradigm (16). The task, described in detail in the next section, measures each leader’s ability to gather information and make decisions. Every lead...
-
[12]
Each problem is self-contained and has an objectively correct solution
Group task: Hidden Profile Each group in the experiment worked collaboratively on a Hidden Profile problem (task materials available in Supplementary Materials). Each problem is self-contained and has an objectively correct solution. The core feature of Hidden Profiles is that...
-
[13]
For example, as groups communicate via text-based chat and the task has a time limit, leaders who are skilled a typing likely have an advantage
Individual assessments 3.1 Measuring ‘Hard Skills’ The group hidden profile task requires leaders and followers to have a set of concrete, hard skills. For example, as groups communicate via text-based chat and the task has a time limit, leaders who are skilled a typing likely...
-
[14]
To focus specifically on the contribution that leaders make independent of their hard skills we condition group performance on measures of leaders hard skills
Identification of leadership skills 4.1 ‘Ground truth’ To identify the total causal contribution that leaders make in the human-only assessment, we exploit the random assignment of leaders to groups. To focus specifically on the contribution that leaders make independent of th...
-
[15]
Bloom, B
N. Bloom, B. Eifert, A. Mahajan, D. McKenzie, J. Roberts, Does management matter? Evidence from India. The Quarterly journal of economics 128, 1–51 (2013)
2013
-
[16]
E. P. Lazear, K. L. Shaw, C. T. Stanton, The Value of Bosses. Journal of Labor Economics 33, 823–861 (2015)
2015
-
[17]
Bloom, R
N. Bloom, R. Sadun, J. Van Reenen, Management as a Technology? (National Bureau of Economic Research Cambridge, MA, 2016; https://www.aeaweb.org/conference/2015/retrieve.php?pdfid=1889&tk=iRZ7HB93)vol. 22327
2016
-
[18]
Managers and productivity in retail
R. D. Metcalfe, A. B. Sollaci, C. Syverson, “Managers and productivity in retail” (National Bureau of Economic Research, 2023); https://www.nber.org/papers/w31192
2023
-
[19]
Weidmann, D
B. Weidmann, D. J. Deming, Team Players: How Social Skills Improve Team Performance. ECTA 89, 2637–2657 (2021)
2021
-
[20]
How Do You Find a Good Manager?
B. Weidmann, J. Vecci, F. Said, D. J. Deming, S. R. Bhalotra, “How Do You Find a Good Manager?” (National Bureau of Economic Research, 2024); https://www.nber.org/papers/w32699
2024
-
[21]
Using large language models to simulate multiple humans and replicate human subject studies
G. V. Aher, R. I. Arriaga, A. T. Kalai, “Using large language models to simulate multiple humans and replicate human subject studies” in International Conference on Machine Learning (PMLR, 2023; https://proceedings.mlr.press/v202/aher23a.html), pp. 337–371
2023
-
[22]
L. P. Argyle, E. C. Busby, N. Fulda, J. R. Gubler, C. Rytting, D. Wingate, Out of one, many: Using language models to simulate human samples. Political Analysis 31, 337–351 (2023)
2023
-
[23]
Dillion, N
D. Dillion, N. Tandon, Y. Gu, K. Gray, Can AI language models replace human participants? Trends in Cognitive Sciences 27, 597–600 (2023)
2023
-
[24]
Large language models as simulated economic agents: What can we learn from homo silicus?
J. J. Horton, “Large language models as simulated economic agents: What can we learn from homo silicus?” (National Bureau of Economic Research, 2023); https://www.nber.org/papers/w31122
2023
-
[25]
J. W. Burton, E. Lopez-Lopez, S. Hechtlinger, Z. Rahwan, S. Aeschbach, M. A. Bakker, J. A. Becker, A. Berditchevskaia, J. Berger, L. Brinkmann, How large language models can reshape collective intelligence. Nature Human Behaviour, 1–13 (2024)
2024
-
[26]
Tranchero, C.-F
M. Tranchero, C.-F. Brenninkmeijer, A. Murugan, A. Nagaraj, THEORIZING WITH LARGE LANGUAGE MODELS. (2024)
2024
-
[27]
Stasser, W
G. Stasser, W. Titus, Pooling of unshared information in group decision making: Biased information sampling during discussion. Journal of personality and social psychology 48, 1467 (1985). 16
1985
-
[28]
J. R. Mesmer-Magnus, L. A. DeChurch, Information sharing and team performance: a meta-analysis. Journal of applied psychology 94, 535 (2009)
2009
-
[29]
L. Lu, Y. C. Yuan, P. L. McLeod, Twenty-Five Years of Hidden Profiles in Group Decision Making: A Meta-Analysis. Pers Soc Psychol Rev 16, 54–75 (2012)
2012
-
[30]
S. G. Sohrab, M. J. Waller, S. Kaplan, Exploring the Hidden-Profile Paradigm: A Literature Review and Analysis. Small Group Research 46, 489–535 (2015)
2015
-
[31]
M. Tomasello, Becoming Human: A Theory of Ontogeny (Harvard University Press, 2019; https://books.google.co.uk/books?hl=en&lr=&id=ZnhyDwAAQBAJ&oi=fnd&pg=PP1&dq =tomasello+becoming+human&ots=5hVNJfakI9&sig=DdxumSUT3- UDhXo8pxHMZr9fzVA)
2019
-
[32]
Raven Progressive Matrices
John, J. Raven, “Raven Progressive Matrices” in Handbook of Nonverbal Assessment, R. S. McCallum, Ed. (Springer US, Boston, MA, 2003; https://doi.org/10.1007/978-1-4615- 0153-4_11), pp. 223–237
2003 doi
-
[33]
Weidmann, Y
B. Weidmann, Y. Xu, PAGE: A Modern Measure of Emotion Perception for Teamwork and Management Research. arXiv arXiv:2410.03704 [Preprint] (2024). http://arxiv.org/abs/2410.03704
2024 arXiv
-
[34]
Stasser, W
G. Stasser, W. Titus, Hidden Profiles: A Brief History. Psychological Inquiry 14, 304–313 (2003)
2003
-
[35]
J. R. Larson, C. Christensen, A. S. Abbott, T. M. Franz, Diagnosing groups: Charting the flow of information in medical decision-making teams. Journal of Personality and Social Psychology 71, 315–330 (1996)
1996
-
[36]
N. K. Steffens, S. A. Haslam, Power through ‘us’: Leaders’ use of we-referencing language predicts election victory. PloS one 8, e77952 (2013)
2013
-
[37]
Weiss, M
M. Weiss, M. Kolbe, G. Grote, D. R. Spahn, B. Grande, We can do it! Inclusive leader language promotes voice behavior in multi-professional teams. The Leadership Quarterly 29, 389–402 (2018)
2018
-
[38]
J. W. Burton, E. Lopez-Lopez, S. Hechtlinger, Z. Rahwan, S. Aeschbach, M. A. Bakker, J. A. Becker, A. Berditchevskaia, J. Berger, L. Brinkmann, L. Flek, S. M. Herzog, S. Huang, S. Kapoor, A. Narayanan, A.-M. Nussberger, T. Yasseri, P. Nickl, A. Almaatouq, U. Hahn, R. H. J. M. ...
2024
-
[39]
J. J. Horton, Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? arXiv arXiv:2301.07543 [Preprint] (2023). http://arxiv.org/abs/2301.07543
2023
-
[40]
Torres, Most people don’t want to be managers
N. Torres, Most people don’t want to be managers. Harvard Business Review 18, 2–4 (2014). 17
2014
-
[41]
Haegele, The Broken Rung: Gender and the Leadership Gap
I. Haegele, The Broken Rung: Gender and the Leadership Gap. arXiv arXiv:2404.07750 [Preprint] (2024). http://arxiv.org/abs/2404.07750
2024 arXiv
-
[42]
D. A. Moore, P. J. Healy, The trouble with overconfidence. Psychological review 115, 502 (2008)
2008
-
[43]
Tsvetkova, T
M. Tsvetkova, T. Yasseri, N. Pescetelli, T. Werner, A new sociology of humans and machines. Nat Hum Behav 8, 1864–1876 (2024)
2024
-
[44]
A. H. Eagly, S. J. Karau, M. G. Makhijani, Gender and the effectiveness of leaders: a meta- analysis. Psychological bulletin 117, 125 (1995)
1995
-
[45]
M. J. Morgan, Women in a Man’s World: Gender Differences in Leadership at the Military Academy. J Applied Social Pyschol 34, 2482–2502 (2004)
2004
-
[46]
A. N. Gipson, D. L. Pfaff, D. B. Mendelsohn, L. T. Catenacci, W. W. Burke, Women and Leadership: Selection, Development, Leadership Style, and Performance. The Journal of Applied Behavioral Science 53, 32–65 (2017)
2017
-
[47]
https://www.economist.com/business/2024/10/03/what-makes-a-good-manager
What makes a good manager? , What makes a good manager?, The Economist. https://www.economist.com/business/2024/10/03/what-makes-a-good-manager
2024
-
[48]
Martin, D
R. Martin, D. J. Hughes, O. Epitropaki, G. Thomas, In pursuit of causality in leadership training research: A review and pragmatic recommendations. The Leadership Quarterly 32, 101375 (2021)
2021
-
[49]
D. J. Deming, The Growing Importance of Social Skills in the Labor Market*. The Quarterly Journal of Economics 132, 1593–1640 (2017)
2017
-
[50]
Economic Decision-Making Skill Predicts Income in Two Countries
A. Caplin, D. Deming, S. Leth-Petersen, B. Weidmann, “Economic Decision-Making Skill Predicts Income in Two Countries” (w31674, National Bureau of Economic Research, Cambridge, MA, 2023); https://doi.org/10.3386/w31674
2023 doi
-
[51]
Westby, C
S. Westby, C. Riedl, Collective Intelligence in Human-AI Teams: A Bayesian Theory of Mind Approach. arXiv arXiv:2208.11660 [Preprint] (2023). https://doi.org/10.48550/arXiv.2208.11660. 18 Table 1. Communication predictors of leader performance. The table presents 12 regression...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.