{"id":"73e6b877-bb01-45d3-b186-73e1e3ea3c0f","arxiv_id":"2607.02198","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Analysis of 53 human-AI team papers yields five distinct clusters (AI Assistant, Ad-hoc Dependency, Ad-hoc Forced Dependency, Paired Equanimity, Group Equanimity) based on psychological team characteristics.","lead":"This paper reviews 53 studies on human-AI teams and sorts them into five clusters drawn from psychological team taxonomies. A smart generalist might read it to see whether findings from one human-AI study can be applied to another or whether the field is studying genuinely different kinds of teams.","discovery_kind":"review","skeptic_critique":{"model":"grok-4.3","headline":"Direct application of human teaming taxonomies to human-AI papers lacks shown validation or adaptation","rationale":"The reader's weakest_assumption matches the load-bearing step exactly. The UNVERDICTED status is appropriate given the absence of mapping details; the concern is internal to the argument rather than external consensus and would be settled by the concrete_test above.","tokens_in":1692,"tokens_out":336,"duration_ms":18133,"concrete_test":"In the methods/results, locate the section describing taxonomy application to the 53 papers; extract the coding scheme or example assignments for at least two clusters and recompute cluster membership after perturbing one taxonomy dimension (e.g., redefining 'dependency' to include AI capability uncertainty); if >20% of papers shift clusters, the five-type structure is sensitive to the unvalidated mapping.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that 53 papers form five distinct clusters (AI Assistant, Ad-hoc Dependency, Ad-hoc Forced Dependency, Paired Equanimity, Group Equanimity) representing disparate team types—rests on categorizing papers 'based on psychological taxonomies of teaming' (abstract). For the clusters to indicate non-transferable insights, the taxonomies must map without substantial distortion. No details are given on how constructs such as 'equanimity' or 'dependency' are operationalized for AI agents (which lack human cognition, shared mental models, or reciprocity), nor on inter-rater reliability, sensitivity to alternative taxonomies, or empirical checks that the resulting clusters differ on measurable team-level variables beyond author judgment.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper reviews 53 studies on human-AI teaming and categorizes them into five clusters (AI Assistant, Ad-hoc Dependency, Ad-hoc Forced Dependency, Paired Equanimity, Group Equanimity) derived from psychological taxonomies of human teaming. It claims these clusters capture distinct holistic team-level characteristics, questions the transferability of insights across the literature, and concludes with reporting guidance and a checklist.","tokens_in":1858,"tokens_out":512,"duration_ms":14299,"significance":"A validated taxonomy that distinguishes team types in human-AI research could aid synthesis and improve comparability across studies; the checklist component offers a concrete contribution if the underlying clusters prove reproducible.","major_comments":[{"comment":"Abstract and Methods: The central categorization into five clusters is described as 'based on psychological taxonomies of teaming' with no reported inclusion/exclusion criteria for the 53 papers, no inter-rater reliability statistics, and no sensitivity analysis to alternative coding schemes or taxonomies; this directly undermines the claim that the clusters represent 'disparate team types' rather than post-hoc groupings.","section":"Abstract / Methods"},{"comment":"Abstract: The mapping of constructs such as 'equanimity' and 'dependency' from human psychological taxonomies to AI agents is asserted without shown operationalization, adaptation, or empirical check that the resulting clusters differ on any measurable team-level variable beyond author judgment.","section":"Abstract"},{"comment":"Results / Discussion: No table or section presents the distribution of papers across clusters, example codings, or quantitative evidence (e.g., cluster separation metrics) that the five types are distinct rather than overlapping or arbitrary.","section":"Results / Discussion"}],"minor_comments":[{"comment":"The abstract states 'we analyse 53 papers' but provides no citation list or supplementary table identifying the papers; this should be added for reproducibility.","section":"Abstract"},{"comment":"Terminology such as 'holistic team-level characteristics' is used without a precise definition or reference to the source taxonomies.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a literature review whose core contribution is the proposed taxonomy; its fit for a methods-oriented HCI venue would be strengthened by explicit validation data or a clear statement that the clusters are exploratory."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments, which highlight opportunities to improve the transparency of our methods and presentation. We address each major comment below and indicate where revisions will be made.","responses":[{"response":"We agree that explicit details on paper selection strengthen the work. The 53 papers were identified via a systematic search using terms such as 'human-AI teaming' and 'human-AI collaboration' in relevant databases, with inclusion focused on studies describing team-level interactions. We will add a Methods section detailing the search strategy and inclusion/exclusion criteria. The categorization followed an iterative, consensus-based process among authors grounded in the cited psychological taxonomies rather than independent multi-rater coding, so inter-rater reliability metrics do not apply; we will note this rationale. Sensitivity analyses to alternative taxonomies lie outside the current scope but can be listed as a limitation. The clusters are not post-hoc but derive directly from combinations specified in the source taxonomies.","revision_made":"partial","referee_comment":"[Abstract / Methods] Abstract and Methods: The central categorization into five clusters is described as 'based on psychological taxonomies of teaming' with no reported inclusion/exclusion criteria for the 53 papers, no inter-rater reliability statistics, and no sensitivity analysis to alternative coding schemes or taxonomies; this directly undermines the claim that the clusters represent 'disparate team types' rather than post-hoc groupings."},{"response":"The constructs are adapted conceptually from established human-team taxonomies, with 'dependency' referring to the extent the AI is required for task completion as described in each paper and 'equanimity' referring to balanced mutual influence. We will add an explicit section or table defining these adaptations and how they were applied to the reviewed papers. No new empirical measurements of team-level variables are provided, as this is a review synthesizing existing literature rather than a primary study; the distinctions rest on the differing profiles of characteristics reported across the papers.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The mapping of constructs such as 'equanimity' and 'dependency' from human psychological taxonomies to AI agents is asserted without shown operationalization, adaptation, or empirical check that the resulting clusters differ on any measurable team-level variable beyond author judgment."},{"response":"We agree that a summary table would improve clarity and will add one in the Results section listing the number of papers per cluster together with representative examples and the defining characteristics for each. Quantitative cluster-separation metrics (e.g., silhouette scores) are not applicable, as the categorization is a theory-driven qualitative mapping to psychological frameworks rather than algorithmic clustering of numerical data. Distinctness is argued via the non-overlapping combinations of team characteristics drawn from the taxonomies.","revision_made":"partial","referee_comment":"[Results / Discussion] Results / Discussion: No table or section presents the distribution of papers across clusters, example codings, or quantitative evidence (e.g., cluster separation metrics) that the five types are distinct rather than overlapping or arbitrary."}],"tokens_in":1313,"tokens_out":656,"duration_ms":30989,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point here is a synthesis that puts 53 papers on human-AI teams into five groups—AI Assistant, Ad-hoc Dependency, Ad-hoc Forced Dependency, Paired Equanimity, and Group Equanimity—using existing psychological team taxonomies. This is new as an organizing frame for the subfield.\n\nIt does a clear job of flagging that studies under the broad label “human-AI team” are not all studying the same setup, which could help stop loose cross-paper claims. The checklist for reporting team type and the closing suggestions for further synthesis are straightforward and usable.\n\nThe soft spot is the direct transfer of human team constructs. The abstract gives no account of how “equanimity” or “dependency” were redefined for an AI partner that lacks shared mental models or reciprocity, nor any inter-rater numbers, inclusion rules, or tests of whether the clusters shift under different coding choices. Without those steps the five groups rest on author judgment alone.\n\nThis is for researchers already working on human-AI collaboration who want tighter language for the setups they study. A reader who needs a practical taxonomy rather than a first-principles theory will find value.\n\nSend it to peer review so the methods and any sensitivity checks can be examined; the core idea is worth testing even if the current evidence for the mapping is thin.","headline":"Paper sorts 53 human-AI team studies into five clusters drawn from human psychology taxonomies, but the mapping step has no visible checks.","tokens_in":2303,"tokens_out":349,"would_cite":false,"duration_ms":22453,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Human-AI teams studied in research fall into five distinct types drawn from psychological taxonomies.","keywords":["human-AI teaming","team types","psychological taxonomies","categorization","human-AI collaboration","team clusters","research synthesis"],"falsifier":"A follow-up review or empirical test that finds most human-AI team studies exhibit characteristics that cut across the five clusters or that the clusters fail to separate the papers meaningfully.","tokens_in":2591,"feed_emoji":"👥","tokens_out":571,"duration_ms":20052,"temperature":0.7,"pith_summary":"The paper reviews 53 studies and sorts them into five clusters using psychological models of human teaming: AI Assistant, Ad-hoc Dependency, Ad-hoc Forced Dependency, Paired Equanimity, and Group Equanimity. Each cluster shows a unique mix of team-level traits, so the single label human-AI teaming covers several different arrangements. A sympathetic reader would conclude that findings from one cluster cannot be assumed to hold for the others. The authors supply a reporting checklist and guidance to make these distinctions clearer in future work.","feed_headline":"Human-AI teams split into five types from psychology models","feed_subtitle":"Review of 53 papers shows disparate structures studied under one label, so insights may not transfer between studies.","key_machinery":"The five-cluster categorization of papers based on psychological taxonomies of teaming, where each cluster is identified by its particular set of team-level characteristics.","core_discovery":"Analysis of the literature shows five separate types of human-AI teams, each defined by a distinct combination of holistic team-level characteristics taken from psychological taxonomies, indicating that the term human-AI team covers multiple disparate structures.","pith_inferences":["Empirical studies could check whether observed human-AI interactions actually align with these five psychological categories.","The mapping may point toward the value of developing team taxonomies built specifically around AI capabilities rather than borrowed human ones.","AI system design choices could be made differently depending on which of the five team types is intended."],"forward_implications":["Insights drawn from studies of one cluster cannot be assumed to transfer to the other four.","Papers must specify which cluster their human-AI team belongs to for findings to be interpretable.","A standardized checklist can help authors report the relevant team-level traits consistently.","Synthesis across the field requires separating the clusters rather than pooling all studies together."],"fun_headline_variants":["Five types of human-AI teams identified in study","53 papers divide human-AI teams into five clusters","Psychology taxonomies reveal five human-AI team types","Human-AI teams grouped into five distinct structures"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That psychological taxonomies of human teaming can be applied directly to human-AI teams without substantial modification.","fun_headline_variants_meta":{"raw":{"variants":["Five types of human-AI teams identified in study","53 papers divide human-AI teams into five clusters","Psychology taxonomies reveal five human-AI team types","Human-AI teams grouped into five distinct structures"]},"model":"grok-4.3","cost_usd":0.004955,"raw_usage":{"total_tokens":2376,"prompt_tokens":573,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":49549500,"prompt_tokens_details":{"text_tokens":573,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1743,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":573,"tokens_out":60,"duration_ms":15510,"temperature":1.0,"reasoning_tokens":1743,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-03T06:02:32.477417+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A follow-up review or empirical test that finds most human-AI team studies exhibit characteristics that cut across the five clusters or that the clusters fail to separate the papers meaningfully.","supporting_citations":[],"review_version":1}