{"id":"07ea09d1-ec4f-496c-b00c-10607bb55b11","arxiv_id":"2508.02966","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a pre-registered lab experiment, leadership performance with AI agents (GPT-4o followers) correlated at 0.81 with a leader's causal impact on human teams, indicating AI agents can proxy for human followers in leadership measurement.","lead":"In a pre-registered experiment, 249 human leaders each led both human teams and teams of AI agents on problem-solving puzzles. A leader's score with the AI agents strongly predicted their measured causal impact on human teams, suggesting cheap AI-based tests might measure leadership skill.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AI follower compliance may inflate the AI–human correlation by reducing the AI test to an individual information-synthesis task, not a test of social leadership.","rationale":"The reader's weakest_assumption—that GPT-4o follower behavior captures human teamwork mechanisms—is exactly the load-bearing point for the paper's headline claim. The paper's own limitation section concedes that emotion plays a diminished role, which is direct evidence that the assumption is only partially met. My concern sharpens this: beyond affect, the AI followers' compliance may remove the social friction that constitutes leadership skill. This is not an internal inconsistency; the statistical analysis is careful (random assignment, pre-registration, multilevel correction for measurement error). It is a construct-validity risk. The existing data can test it: the OSF repository contains chat logs and code, so the proposed behavioral coding is feasible without new data collection. Because the reader already assigned CONDITIONAL, and this concern is a specified condition (validate follower behavior) rather than a demonstrated fatal flaw, the verdict should remain CONDITIONAL. The paper deserves credit for pre-registration, random assignment, and posted materials; the concern is about the interpretation of the strong correlation, not about the internal validity of the experiment.","tokens_in":11421,"tokens_out":5847,"duration_ms":72375,"concrete_test":"Use the existing chat logs to compare follower behavior between the human and AI conditions. Code each utterance where a follower holds a private clue, and estimate: (i) P(private clue disclosed | leader asked) versus P(disclosed | not asked); (ii) rate of volunteered private clues; (iii) rate of misinterpretation or irrelevant responses. If the AI condition shows near-ceiling disclosure conditional on being asked, and significantly lower misinterpretation than the human condition, the AI test removes the social friction that leadership skill is supposed to overcome. A complementary check: re-estimate the disattenuated correlation between AI and human leader scores after residualizing both on a composite of all measured cognitive skills (task skill, fluid IQ, reading/typing speed); if the residual correlation drops below roughly 0.4, the proxy claim is substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central estimate (ρ̂=0.81, Section 3) is interpreted as evidence that LLM agents are effective proxies for human participants. That interpretation requires GPT-4o followers to reproduce the follower behaviors that make leader skill causally matter in human teams. The paper validates this only at the level of leader outcomes (similar predictor correlations, similar leader effects), not at the level of follower behavior. The authors openly note that affect is diminished in the AI test, but a more damaging possibility is compliance: if AI followers disclose all unshared clues whenever directly asked, never misinterpret or withhold information, and never react emotionally, then the AI test mostly measures the leader's question-asking and integration—cognitive skills already partly captured by the individual hidden-profile task and fluid IQ. In human teams, leaders must also overcome social friction (reluctance, distraction, misinterpretation), which is exactly the 'soft skill' component the paper claims to measure. The raw correlation of 0.67 and disattenuated 0.81 could then reflect common cognitive ability rather than a genuine proxy for social leadership. Conditioning on measured hard skills (Eq. 1–2) does not eliminate this concern, because unmeasured verbal/executive abilities could still load on both tests. Table 1 shows that leaders' question-asking predicts success in both conditions, but it does not show that followers respond to questions in the same way. Thus the claim that the AI test captures the causal mechanisms of human teamwork is load-bearing and unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a pre-registered lab experiment (n=249 leaders) in which each leader directed teams of three human followers and, separately, teams of three GPT-4o AI followers on a modified Hidden-Profile task. Leaders were randomly assigned to multiple human teams to estimate their causal contribution to group performance, and this ground-truth measure was compared with the same leaders' performance on an AI version of the task. The authors report a raw leader-level correlation of 0.67 between AI and human test performance, rising to a disattenuated correlation of 0.81, and similar patterns of predictors (skill measures, demographics) and communication behaviors across the two tests. They conclude that AI agents can serve as effective proxies for human participants in social experiments, enabling scalable measurement of leadership skills.","tokens_in":11670,"tokens_out":8475,"duration_ms":100633,"significance":"If the central correlation is robust and generalizable, the paper offers a low-cost, performance-based method for measuring leadership skill, with implications for personnel selection, leadership training, and experimental social science. The design has notable strengths: it is pre-registered, the order of AI and human tests is counterbalanced, leaders are repeatedly randomly assigned to human teams, hard-skill confounders (typing speed, individual hidden-profile skill, fluid IQ) are measured and conditioned on, process measures (question asking, turn-taking, pronoun use) are extracted from chat logs, and the data, code, and materials are promised publicly. The reported correlations are large and the predictor profiles are similar across tests. However, the headline disattenuated correlation is not fully transparent, the equivalence of the two puzzle forms is asserted rather than demonstrated, and the 'AI agents as proxies' claim goes beyond the leader-level evidence provided.","major_comments":[{"comment":"The headline estimate ρ̂=0.81 is a disattenuated correlation, but the manuscript never specifies the reliability coefficients or the attenuation-correction formula used. Without this information the reader cannot judge whether 0.81 is an over-correction, which is load-bearing for the paper's main claim. Please report the reliability estimates (e.g., split-half or intraclass correlations) for Y_i^AI and Y_i^Human, the exact correction formula, and a sensitivity analysis using alternative reliability values. Also, the confidence interval is printed as '[0.72,88]' and should read '[0.72,0.88]'.","section":"Section 3, 'Individual scores on the AI test correlate very highly with ground-truth scores'"},{"comment":"The claim that the two parallel puzzle forms have 'equivalent structure and difficulty across items' is asserted without supporting evidence. If the human and AI forms differ in difficulty or in the distribution of clue types, the leader-level correlation between the two tests could be inflated or attenuated. Because every leader took both forms, a form-order confound would directly affect the cross-test correlation. Please provide evidence of form equivalence (e.g., pilot calibration data or pretest statistics) and clarify whether the forms were randomly assigned or counterbalanced with respect to the human/AI test order and the two-order condition.","section":"Section 2, 'Group task'"},{"comment":"The conclusion that 'AI agents can be effective proxies for human participants in social experiments' goes beyond the evidence presented. The study validates the AI test at the level of leader outcomes (high correlation with human-team leadership, similar predictor profiles), but it does not demonstrate that AI followers behave like human followers: the chat logs are not analyzed to compare information disclosure, misinterpretation, compliance, or emotional reactions. The authors' own results show that positive affect does not predict AI-team performance and that emotional perceptiveness is a weaker predictor of AI-test success, which indicates that the AI followers are not fully human-like. To support the proxy claim, either the conclusions should be scoped to 'AI-based assessment captures leader skills that transfer to human teams,' or the authors should add follower-level behavioral comparisons from the chat data. Without such evidence, the 0.81 correlation could partly reflect common cognitive skill (e.g., verbal ability or executive function) rather than social leadership, and conditioning on measured hard skills (Eqs. 1-2) does not eliminate this concern because unmeasured abilities could still load on both tests.","section":"Abstract and Discussion"}],"minor_comments":[{"comment":"The subscripts in equation (1) are inconsistent: the sum is written as Σ_g I_{ig} X_i, but the left-hand side is indexed by g; this should likely be Σ_i I_{ig} X_i, with the residual indexed by g (ε_g) rather than ε_{ig}. Please clarify the notation so the definitions of α_i in equation (2) are unambiguous.","section":"Methods, Eqs. (1)-(2)"},{"comment":"The sentence 'mirroring findings from the field (8)' cites reference 8 (Argyle et al., on simulating human samples), but the intended reference is presumably to the empirical literature on gender, ethnicity, and leadership (e.g., references 30–32). Please correct this citation.","section":"Discussion"},{"comment":"The caption describes the left panel as showing 'Leadership Skill,' whereas the text says the left panel conditions on task-specific skill. Please make the caption consistent with the text (e.g., 'leadership skill after conditioning on hard skills').","section":"Figure 5 caption"},{"comment":"The term 'disattenuated correlation' is used without a parenthetical explanation. Consider adding a brief note defining the correction for measurement error, as many readers will not be familiar with psychometric terminology.","section":"Abstract and main text"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically strong and likely to attract wide interest. The main risks are the lack of transparency in the disattenuation procedure and the unverified equivalence of the two puzzle forms. The authors should be asked to provide the reliability details and form-equivalence evidence; the broader 'effective proxies' claim should be either supported with follower-level behavioral analyses or appropriately scoped. I see no reason to reject the paper, but these points are load-bearing and need to be addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Best quick take: this is a serious, pre-registered measurement study. The authors estimate causal leader effects in human teams by randomly assigning 249 leaders to multiple groups, then show that a cheaper AI-agent version of the same task predicts those leader effects well (disattenuated rho = 0.81, raw 0.67). That is new and worth taking seriously. The design is careful: counterbalanced order, six groups per leader, public data and code, and conditioning on hard skills gives a softer but still substantial 0.69 correlation.\n\nThe soft spots are real but mostly addressable. The biggest one is exactly what the stress-test note flags. The paper interprets the correlation as evidence that LLM agents can stand in for human followers, but it only validates outcome-level similarity, not the behavior that generates it. If GPT-4o followers are overly compliant—always disclosing clues when asked, never withholding, misreading, or getting distracted—then the AI test becomes mostly an individual information-synthesis task. The cognitive-loading on fluid IQ and verbal ability could explain much of the correlation, and conditioning on two measured hard skills doesn't eliminate that because unmeasured executive or verbal abilities could load on both tests. The authors admit emotion plays a smaller role in the AI condition, which is a direct sign the follower behavior differs. They don't show the conversational dynamics are similar beyond coarse counts of questions and turn-taking. So the proxy claim is load-bearing and unverified.\n\nOther soft spots: the disattenuated rho needs more transparency about the reliability corrections; the two puzzle forms are asserted equivalent but never checked; and the abstract's claim that AI agents are effective proxies for human participants in social experiments goes well beyond one hidden-profile task. The model is also an unversioned closed GPT-4o, which makes replication fragile.\n\nNone of this is fatal. The empirical result is what it is: a strong correlation between a cheap AI test and a human ground truth, in a single task. The paper would be more convincing with follower-level analysis, reliability details, and a more modest abstract. I'd send it to review, and I'd cite it. This is exactly the kind of work a serious referee should engage with.","headline":"A serious, pre-registered measurement study with a real and interesting correlation, but the abstract overclaims that AI agents are effective proxies for human participants without validating follower behavior.","tokens_in":12196,"tokens_out":2799,"would_cite":true,"duration_ms":32999,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new experiment claims that a leader's skill with AI agents predicts their causal impact on human teams, with a disattenuated correlation of 0.81.","keywords":["leadership measurement","AI leadership test","hidden profile","large language models","causal contribution","teamwork skills","silicon samples","pre-registered experiment"],"falsifier":"Re-run the same pre-registered design with a different large language model or with follower prompts that reduce cooperativeness or information sharing; if the disattenuated correlation with human ground-truth causal contributions drops well below 0.81, or if the leader-effect sizes diverge, the proxy is specific to the original agent configuration rather than to general leadership skill.","tokens_in":11223,"feed_emoji":"🤖","tokens_out":3683,"duration_ms":43399,"temperature":0.7,"pith_summary":"This paper asks whether the ability to lead groups of humans can be measured by leading groups of AI agents. In a pre-registered lab study, 249 human leaders each solved hidden-profile puzzles with both human and AI followers, and the researchers estimated each leader's causal contribution to team performance through repeated random assignment. The central claim is that performance on the AI leadership test is very highly correlated with the ground-truth human measure, especially after correcting for measurement error (disattenuated rho = 0.81). If correct, this would mean that large language model agents can serve as scalable, low-cost proxies for human participants in social experiments on leadership and teamwork.","feed_headline":"AI agents can score leadership skill that works on humans","feed_subtitle":"In 249 leaders, AI-follower scores correlated at 0.81 with causal impact on human teams, after correcting for measurement error.","key_machinery":"The central object is the AI leadership test built on a modified Hidden Profile paradigm: a leader and three followers each hold private clues, the leader can talk to anyone but followers can only talk to the leader, and the leader must synthesize dispersed information into probabilistic answers. The ground-truth measure is the leader's average causal contribution estimated by randomly assigning each leader to multiple human groups and then using a multilevel model to isolate the leader-effect standard deviation. The paper's key comparison is between leader-effect estimates from the AI version and the human version, with parallel puzzle forms created so that the AI agents would not have seen solutions in training data.","core_discovery":"The paper's central discovery is that individual scores on an AI leadership test, where a human directs three LLM followers through collaborative puzzles, strongly predict the same leader's total causal contribution to teams of human followers. Using a hidden-profile task with a distinct leader role and probabilistic answers, the authors randomly assigned each leader to six human groups and six AI-agent groups. The disattenuated correlation between AI-test scores and human ground-truth scores is rho = 0.81 (95% CI [0.72, 0.88]); after conditioning on leader hard skills such as fluid IQ, typing speed, and task-specific ability, the correlation remains 0.69 (95% CI [0.57, 0.81]). The same leader characteristics predict success in both settings, and the same communication behaviors (asking questions, conversational turn-taking, using 'we' language) are associated with performance. The authors interpret this as evidence that LLM agents can be effective proxies for human participants in measuring leadership and teamwork skills.","pith_inferences":["If the proxy relationship holds, the same repeated-random-assignment design could measure causal leader contributions longitudinally, enabling pre/post evaluation of leadership training programs that currently lack robust outcome measures.","The paper's own evidence that positive affect matters less in the AI test suggests that AI-follower assessments may underweight socioemotional leadership; a test battery that deliberately varies follower affect or cooperativeness could recover that dimension.","The raw (non-disattenuated) correlation is 0.67, so in practical selection settings with short, noisy assessments the predictive validity will be lower than the headline 0.81; operational use would need multiple sessions or longer tests.","The hidden-profile task is one specific team process; whether the human-AI correlation generalizes to other team activities, such as creative brainstorming or crisis decision-making, remains an open testable question."],"forward_implications":["Leadership assessment becomes scalable: the AI version cost about $23 per leader and ran autonomously, versus $114 and live coordination for the human version, so repeated-random-assignment measurement could be used much more widely.","The AI test predicts not just raw leadership success but also leadership-specific soft skills, since the correlation survives conditioning on hard skills (disattenuated rho = 0.69).","Substantive findings about leadership, such as overconfidence predicting willingness to lead and accurate self-evaluators contributing more, replicate in the AI test, suggesting silicon samples can test leadership hypotheses.","Because demographics do not predict leadership skill, a cheap standardized AI-based test could in principle identify high-potential leaders overlooked by current selection processes."],"supporting_citations":[{"why":"Supplies the repeated-random-assignment method for estimating leaders' causal contributions to team performance, which defines the ground truth.","marker":"(5)"},{"why":"Extends the causal-contribution approach to managers, providing the framework for decomposing leader effects from hard skills.","marker":"(6)"},{"why":"Demonstrates that large language models can simulate multiple humans and replicate human subject studies, the premise for using LLM followers.","marker":"(7)"},{"why":"Shows that language-model samples can approximate human survey samples, supporting the 'silicon sample' approach the paper extends to team settings.","marker":"(8)"},{"why":"Introduces the hidden-profile paradigm in which dispersed information must be pooled through discussion, the basis of the group task.","marker":"(13)"},{"why":"Reviews hidden-profile research and motivates information pooling as a core component of successful teamwork.","marker":"(16)"},{"why":"Provides the PAGE emotion-perception measure used to test whether social intelligence predicts leadership in both AI and human settings.","marker":"(19)"},{"why":"Describes an online hidden-profile implementation with human-AI teams, informing the task interface and probabilistic answer design.","marker":"(37)"}],"fun_headline_variants":["AI agents predict leadership impact on human teams","AI-follower test scores match human leadership success","Your AI leadership score forecasts real team influence","AI agents as proxies reveal who leads humans best"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result stands or falls on whether LLM followers, prompted with instructions analogous to those given to human followers, respond to leader behavior through the same causal channels (information pooling, question asking, turn-taking) that make leaders effective with human teams.","fun_headline_variants_meta":{"raw":{"variants":["AI agents predict leadership impact on human teams","AI-follower test scores match human leadership success","Your AI leadership score forecasts real team influence","AI agents as proxies reveal who leads humans best"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3230,"prompt_tokens":876,"completion_tokens":2354,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":2297}},"tokens_in":492,"tokens_out":2354,"duration_ms":16173,"temperature":1.0,"reasoning_tokens":2297,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:45:36.341566+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same pre-registered design with a different large language model or with follower prompts that reduce cooperativeness or information sharing; if the disattenuated correlation with human ground-truth causal contributions drops well below 0.81, or if the leader-effect sizes diverge, the proxy is specific to the original agent configuration rather than to general leadership skill.","supporting_citations":[],"review_version":1}