REVIEW 4 major objections 5 minor 16 references
Improving the State of the Art for Training Human-AI Teams: Technical Report #5 -- Individual Differences and Team Qualities to Measure in a Human-AI Teaming Testbed
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A technical review selects the individual-trait and team-quality surveys to include in a human-AI teaming testbed.
desk verdict A competent, useful literature review for assembling a testbed survey battery, but it offers no new findings and silently assumes human-human measures transfer to human-AI teams. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a two-tier measurement framework: stable trait-level constructs that people bring into any team, and team-specific beliefs and attitudes that can shift with experience on a particular team. The framework distinguishes individual-level scores from team-level aggregates such as Team Personality Elevation and Team Personality Diversity, and it treats referent-shift constructs like team efficacy as individual perceptions referenced to the team. The suitability criteria (public availability, psychometric validation, applicability to the task environment, and applicability to the target participant) do the work of filtering dozens of candidate instruments down to the final battery, with the Autonomous Agent Teammate-likeness scale serving as the one measure built specifically for rating an AI agent as a teammate.
What would settle it
A confirmatory factor analysis or measurement-invariance study of the recommended scales in a human-AI teaming sample would settle this: if the factor structure of, say, the ten-item Revised Group Environment Questionnaire or the teamwork scales from Campion et al. fails to replicate, or loadings differ materially from human-human reference samples, the suitability judgments in Tables 10 and 11 would be undercut.
Extended reading notes
Core claim
The central claim is that the measures summarized in Tables 10 and 11 are suitable for inclusion in the Concept of Operations for the Synthetic Task Environment. The recommended trait battery consists of the IPIP-NEO-120 and the Mini IPIP for Big Five personality, the Trait Emotional Intelligence Questionnaire for emotional intelligence, the Spheres of Control scales for locus of control, and the Preference for Group Work and Preference for Virtual Teams measures. The recommended team-level battery consists of the social support, communication and cooperation, self-management, workload sharing, and participation scales, the Individual Differences in Anthropomorphism Questionnaire, the ten-item Revised Group Environment Questionnaire for cohesion, and the Autonomous Agent Teammate-likeness scale. The report deliberately favors publicly available, self-report instruments with broad item wording, because the testbed needs measures that can be administered before and after a team performance scenario and that will work for participants who may be meeting both human and AI teammates for the first time.
Load-bearing premise
The report assumes that questionnaires validated mainly on human-human teams will still measure the same constructs when one or more teammates is an AI agent, since none of the recommended instruments except the Autonomous Agent Teammate-likeness scale is revalidated in a human-AI setting.
Editorial extensions
If this is right
- Researchers can field the recommended battery as a pre- and post-performance survey package in human-AI teaming experiments.
- Because all recommended measures are publicly available and self-report, other laboratories can replicate the testbed's data collection without licensing costs.
- Trait scores can be aggregated at the team level as elevation or diversity, allowing tests of whether human-team personality composition effects reappear when an AI agent is a teammate.
- Team-specific measures such as cohesion, efficacy, and teammate-likeness can detect change across training or repeated missions.
- The Autonomous Agent Teammate-likeness scale gives a direct outcome measure of whether humans perceive an AI agent as a tool or a teammate, and the anthropomorphism questionnaire gives a trait-level covariate for that perception.
Reading between the lines
- The paper leaves implicit that every recommended measure except the Autonomous Agent Teammate-likeness scale was validated in human-human teams, so the full battery carries an untested assumption that AI teammates do not change what the items measure.
- One testable extension is to reword the Preference for Group Work and Preference for Virtual Teams items to compare working with an AI agent against working alone, directly probing the tool-versus-teammate distinction the report raises.
- The Big Five team-composition effects the report reviews could be re-estimated on human-AI teams using the recommended measures, yielding the first direct evidence on whether those effects transfer when one teammate is an agent.
- Pairing the anthropomorphism questionnaire with the Autonomous Agent Teammate-likeness scale would allow a testbed experiment to manipulate agent form or behavior and test whether changes in perceived teammate-likeness depend on the human's tendency to anthropomorphize.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This technical report, the fifth in a series from Sonalysts, reviews individual-difference and team-quality constructs that the authors propose to include in a Concept of Operations (CONOPS) for a Synthetic Task Environment (STE) for human-AI teaming research. The report surveys Big Five personality, emotional intelligence, individualism/collectivism, locus of control, preference for teamwork, teamwork attitudes, team efficacy, and team cohesion, summarizing for each measure its availability, reliability/validity evidence, report type, and applicability to the task and participant. It concludes with two tables (Tables 10 and 11) listing recommended measures for inclusion in the CONOPS, including the Mini-IPIP, IPIP-NEO-120, TEIQue, Spheres of Control, Preference for Group Work, Preference for Virtual Teams, Campion's teamwork subscales, IDAQ, the 10-item Revised GEQ, and the Autonomous Agent Teammate-likeness (AAT) scale. The central claim is that these measures are suitable for collecting meaningful pre- and post-performance survey data in the proposed human-AI teaming testbed.
Significance. If the suitability judgments are accepted, the report provides a useful, well-organized menu of candidate instruments for researchers building human-AI teaming testbeds. Its strengths include explicit reporting of item availability, concise summaries of psychometric evidence, reproduction of item content for several measures (e.g., Psychological Collectivism, Spheres of Control, T-TAQ, the ten-item GEQ), and a clear separation of trait-level from team-level constructs. The inclusion of the AAT scale and the IDAQ, which are directly relevant to human-AI interaction, is a constructive step beyond generic human-team measures. However, the report's contribution is limited by two load-bearing gaps: it never describes the STE or CONOPS enough to make applicability judgments meaningful, and it assumes without evidence that instruments validated on human-human teams will retain their psychometric properties when a teammate is an AI agent. These gaps affect the central recommendations rather than only the presentation.
major comments (4)
- [§1 and §4] The suitability ratings in Tables 10 and 11 are made relative to 'the task environment' and 'the target participant,' yet the report never specifies the Synthetic Task Environment's tasks, team composition, communication modality, or participant population. For example, Table 9 rates the 10-item Revised GEQ as 'Suitable' while its own applicability note says the items presume co-located teams, and the report elsewhere raises distributed teams as a possibility. Without a CONOPS description, the reader cannot verify whether the 'Applicability to the Task Environment' and 'Applicability to the Target Participant' columns are accurate. Please add a concrete description of the STE, the anticipated human-AI team structure, and the intended participant population, or explicitly frame the tables as conditional on assumptions that are stated.
- [§3.3.1, §4 (Tables 10-11)] The central recommendation assumes construct continuity from human-human to human-AI teams. The report itself notes in §3.3.1 that the AAT scale was developed specifically for human-AI contexts, but all other recommended measures (Big Five inventories, Spheres of Control, Preference for Group Work, Campion's social support/communication/self-management/workload-sharing scales, and the 10-item Revised GEQ) carry validity evidence from human-human teams. No plan is provided to test measurement invariance, factor structure, or criterion-related validity when one or more teammates is an AI agent. Because the stated purpose is collecting data interpretable against the human-teaming literature, this is a load-bearing gap. Please add an explicit validation strategy for the human-AI context, or temper the suitability recommendations to 'candidate measures requiring revalidation.'
- [Table 9, §3.3.1] There is an internal contradiction in the treatment of the 10-item Revised GEQ (Carless & De Paola, 2000). The table's 'Applicability to the Target Participant' row states that the items 'are specific to teams that are co-located, and so would not apply to distributed teams (which may be important for the research conducted in the proposed testbed),' yet the 'Suitability for Inclusion in CONOPS' column concludes 'Suitable.' If the testbed may involve distributed human-AI teams, this measure should be marked as requiring adaptation or as conditionally suitable, not unconditionally suitable. This inconsistency directly affects the final list in Table 11.
- [§3.1.2, Table 6] The report recommends Campion's Social Support and Communication/Cooperation subscales as components of a larger battery, and later includes them in Table 11, but it also notes these subscales were not validated against similar measures and are not generalized attitude measures. More importantly, the items (e.g., 'Members of my team cooperate to get the work done') presuppose human team members with shared intentionality and mutual accountability. The report does not discuss whether these items are meaningful when one team member is an AI agent. Please either provide a rationale for why these items are expected to function in human-AI teams or restrict the recommendation to human-human baseline comparisons.
minor comments (5)
- [§2.3.1] The name 'Mayer and Slovey (1997)' appears in the text; this should be 'Salovey' (Mayer & Salovey, 1997). The same typo appears in the reference list as 'Mayer, J.D., & Salovey, P.' in the correct form, so this is a simple spelling inconsistency.
- [§2.6] The text says 'another sub-dimension of the individualism/collectivism dichotomy discussed in section 2.3.' Individualism/collectivism is covered in Section 2.4, not Section 2.3 (which is Emotional Intelligence). The cross-reference should be corrected to Section 2.4.
- [§3.3.1 and Table 9] The author name is inconsistently spelled as 'Caron' and 'Carron' (e.g., 'Caron et al., 2003' and 'Carron et al., 1985'). The reference list uses 'Carron.' Please standardize to the correct spelling throughout, and also correct 'Zaccarow and Lowe' to 'Zaccaro and Lowe.'
- [Table 4] The 'Availability' entries for the Spheres of Control (Paulhus, 1983) and the Rotter IE scale are inconsistent with the surrounding text: the table says 'Publically Available' for Spheres of Control but the text does not provide a source or link, while the Rotter IE scale is marked 'Not Publically Available' even though the Rotter IE scale is widely reprinted. Please clarify the operational definition of 'publically available' and provide a consistent basis for these judgments.
- [§3.2.1] In the description of Riggs and Knight's (1994) collective efficacy items, the reverse-scored items are marked '(R)' but no scoring instructions are given for the overall scale. Similarly, the report does not state whether the Guzzo et al. (1993) team potency items are summed or averaged. Adding a sentence on scoring would improve reproducibility for readers who want to use the items.
Circularity Check
No circularity: the report is a literature review that makes no derived predictions or fitted claims; its recommendations rest on external published measures and are not equivalent to their own inputs.
full rationale
This technical report is a scoping literature review whose output is a set of suitability judgments (Tables 10 and 11) for survey instruments to include in a future Synthetic Task Environment. It contains no fitted parameters, no equations, and no empirical predictions that could reduce to their own inputs. Each suitability judgment is explicitly grounded in external sources (e.g., Donnellan et al., 2006 for Mini-IPIP; Carless and De Paola, 2000 for the 10-item GEQ; Wynne and Lyons, 2018/2019 for the AAT), and the report repeatedly flags limitations rather than hiding them (e.g., Section 3.3.2 notes that the 10-item GEQ items presume co-located teams; Table 3 notes that Psychological Collectivism may require adaptation; Section 3.3.1 identifies the AAT as the one measure developed specifically for human-AI contexts). The load-bearing assumption that human-human validity evidence transfers to human-AI teams is a substantive external-validity concern, not a circularity: it is an assumption about the world, not a claim that is true by construction or by self-citation. No self-citations by the authors are used as evidence, and no uniqueness theorem or ansatz is imported from prior work to force a conclusion. The report is therefore self-contained as a review, and no circular step can be exhibited.
Assumptions & free parameters
assumptions (2)
- domain assumption Measures validated on human-human teams are assumed to transfer to human-AI teams without revalidation.
- ad hoc to paper The authors' suitability criteria (public availability, reliability/validity, applicability to task and participant) are treated as sufficient for inclusion in an unvalidated CONOPS.
Cite this review
Pith. "Pith review of Improving the State of the Art for Training Human-AI Teams: Technical Report #5 -- Individual Differences and Team Qualities to Measure in a Human-AI Teaming Testbed." pith.science (2026). https://pith.science/paper/GVSNYFOS
@misc{pith2026250718878,
author = {Pith},
title = {Pith review of: Improving the State of the Art for Training Human-AI Teams: Technical Report #5 -- Individual Differences and Team Qualities to Measure in a Human-AI Teaming Testbed},
year = {2026},
howpublished = {\url{https://pith.science/paper/GVSNYFOS}},
note = {Machine review of arXiv:2507.18878}
}
read the original abstract
Sonalysts, Inc. (Sonalysts) is working on an initiative to expand our expertise in teaming to include Human-Artificial Intelligence (AI) teams. The first step of this process is to develop a Synthetic Task Environment (STE) to support our original research. Prior knowledge elicitation efforts within the Human-AI teaming research stakeholder community revealed a desire to support data collection using pre- and post-performance surveys. In this technical report, we review a number of constructs that capture meaningful individual differences and teaming qualities. Additionally, we explore methods of measuring those constructs within the STE.
Reference graph
Works this paper leans on
-
[1]
Baker, D. P., Amodeo, A. M., Krokos, K. J., Slonim, A., & Herrera, H. (2010). Assessing teamwork attitudes in healthcare: development of the TeamSTEPPS teamwork attitudes questionnaire . Quality and Safety in Health Care, 19(6), e49-e49. Bandura, A. (2006). Guide for constructing self-efficacy scales. Self-efficacy beliefs of adolescents, 5(1), 307-337. B...
work page 2010
-
[2]
Waytz, A., Cacioppo, J., & Epley , N. (2010). Who sees human? The stability and importance of individual differences in anthropomorphism. Perspectives on Psychological Science, 5(3), 219-232. Wong, C. S., & Law, K. S. (2002). Wong and law emotional intelligence scale. The leadership quarterly. Wynne, K. T., & Lyons, J. B. (2018). An integrative model of a...
work page 2010
-
[27]
Shaw, J. D., Duffy, M. K., & Stark, E. M. (2000). Interdependence and preference for group work: Main and congruence effects on the satisfaction and performance of group members . Journal of Management, 26(2), 259-279. Spector, P. E. (1988). Development of the work locus of control scale. Journal of occupational psychology, 61(4), 335-340. Spector, P. E.,...
work page 2000
-
[130]
Seal, C. R., Beauchamp, K. L., Miguel, K., & Scott, A. N. (2011). Development of a self‐report instrument to assess social and emotional development . Journal of Psychological Issues in Organizational Culture, 2(2), 82-95. Seal, C . R., Miguel, K., Alzamil, A., Naumann, S . E., Royce -Davis, J., & Drost, D . (2015). Personal- interpersonal competence asse...
work page 2011
-
[192]
Festinger, L. (1950). Informal Social Communication. Psychological review, 57(5),
work page 1950
-
[271]
70 Gibson, C. B. (1999). Do they do what they believe they can? Group efficacy and group effectiveness across tasks and cultures. Academy of Management Journal, 42, 138-152. Goleman, D., Boyatz is, R., & McKee, A . (2002). Primal leadership: Realizing the power of emotional intelligence. Boston: Harvard University Press. Guzzo, R . A., Yost, P . R., Campb...
work page 1999
-
[288]
Carless, S . A., & De Paola, C . (2000). The measurement of cohesion in work teams . Small group research, 31(1), 71-88. Carron, A. V., Brawley, L . R., Eys, M . A., Bray, S., Dorsch, K., Estabrooks, P., .. . & Terry, P . C. (2003). Do individual perceptions of group cohesion reflect shared beliefs? An empirical analysis. Small group research, 34(4), 468-...
work page 2000
-
[376]
Neuman, G . A., Wagner, S . H., & Christiansen, N . D. (1999). The relationship between work -team personality composition and the job performance of teams. Group & Organization Management, 24(1), 28-45. Offermann, L. R., Bailey, J. R., Vasilopoulos, N. L., Seal, C., & Sass, M. (2004). The relative contribution of emotional competence and cognitive abilit...
work page 1999
Show all 16 references
-
[595]
M., Martí-Vilar, M., Merino-Soto, C., & Cervera-Santiago, J
Bru-Luna, L. M., Martí-Vilar, M., Merino-Soto, C., & Cervera-Santiago, J. L. (2021). Emotional intelligence measures: A systematic review. In Healthcare (Vol. 9, No. 12, p. 1696). MDPI. Butler, L., Park, S. K., Vyas, D., Cole, J. D., Haney, J. S., Marrs , J. C., & Williams, E....
2021
-
[755]
H., & Knafo, A
Roccas, S., Sagiv, L., Schwartz, S. H., & Knafo, A. (2002). The big five personality factors and personal values. Personality and social psychology bulletin, 28(6), 789-801. Salovey, P., Mayer, J. D., Goldman, S. L., Turvey, C., & Palfai, T. P. (1995). Emotional attention, cla...
2002
-
[847]
Big Five Questionnaire
Caprara, G. V., Barbaranelli, C., Borgogni, L., & Perugini, M . (1993). The “Big Five Questionnaire”: A new questionnaire to assess the five factor model . Personality and individual Differences, 15(3), 281-
1993
-
[851]
L., Colquitt, J
Jackson, C. L., Colquitt, J . A., Wesson, M . J., & Zapata-Phelan, C. P. (2006). Psychological collectivism: A measurement validation and linkage to group member performance . Journal of applied psychology, 91(4),
2006
-
[884]
Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Mod el with a 120 -item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51, 78-89. Kapalo, K. A., Phillips, E., & Fiore, S. M. (2016, September). The Application and...
2014
-
[989]
A., Rodriguez-Guzman, J., Baladrón-González, V.,
Muñoz de Morales -Romero, L., Bermejo-Cantarero, A., Martínez -Arce, A., González -Pinilla, J . A., Rodriguez-Guzman, J., Baladrón-González, V., ... & Redondo-Calvo, F. J. (2021). Effectiveness of an Educational Intervention With High -Fidelity Clinical Simulation to Improve A...
2021
-
[1253]
A., Rutte, C
Peeters, M. A., Rutte, C. G., van Tuijl, H . F., & Reymen, I . M. (2006a). The big five personality traits and individual satisfaction with the team. Small group research, 37(2), 187-211. Peeters, M. A., Van Tuijl, H. F., Rutte, C. G., & Reymen, I. M. (2006b). Personality and ...
2006
-
[1832]
Mayer, J. D. (2002). MSCEIT: Mayer-Salovey-Caruso emotional intelligence test. Toronto, Canada: Multi- Health Systems. Mayer, J.D., & Salovey, P . (1997). What is emotional intelligence? In P . Salovey & D.J . Sluyter (Eds.), Emotional development and emotional intelligence: E...
2002
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.