{"id":"162af88f-7a88-456e-9b2a-a03c46d67003","arxiv_id":"2607.27179","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"In a moral-dilemma chat RCT, an AI teammate dominated airtime with low-newness talk and immediately reduced human-human responsivity, belonging, and felt value.","lead":"Adding an AI teammate to a three-person chat team made the AI the most talkative member while cutting how much the two humans responded to each other, and they felt less belonging and status. The finding matters because product teams are shipping AI as a peer, not a tool, often without measuring what that does to human-human ties.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Group-size and GCA-node asymmetry still undercuts causal attribution of the human-human social cost to AI presence.","rationale":"The reader correctly isolates the design confound as the soft spot under the strongest claim. RQ1 (within-team AI role) and the within-treatment airtime–value slope are informative and less threatened. RQ2’s causal story and the “immediate baseline” reading of RQ3 nulls are not. No stronger internal inconsistency appeared: methods are standard GCA/LIWC, direction of effects is coherent, and surveys partially corroborate discourse. That still leaves CONDITIONAL as the right call—build on after size-matched controls and clearer human-only edge definitions—not REJECT. I do not elevate secondary issues (Bonferroni-borderline belonging, unreleased prompts/code, gender imbalance) above the group-size/network issue, because fixing the latter is what would most change confidence in the headline social-cost claim.","tokens_in":13716,"tokens_out":579,"duration_ms":29174,"concrete_test":"Recompute treatment GCA responsivity/social impact on student–student uptake only (drop the AI node), and compare those dyadic rates to mean pairwise human–human rates inside control triads (same window/embedding settings). In parallel, or as the decisive design fix, collect a 2-human no-AI arm. If the treatment deficit on responsivity, social impact, belonging, and status shrinks to non-significance under either check, the social-cost attribution to AI presence does not hold as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central between-condition claim (RQ2) attributes lower human-human responsivity, social impact, belonging, and status to AI presence. The design substitutes AI for a human (2 students+AI vs 3 students), so human count and discourse-network composition are confounded with condition (§3.1, §3.5). Per-student scoring and within-pipeline z-standardization do not remove that structure: treatment humans have one human partner plus an AI node in the GCA graph, while controls have two human partners and no AI node. Survey composites are independent of GCA but still compare dyads-with-AI to human triads, which differ in baseline relational dynamics (Wheelan 2009, cited). The paper’s own defense—that smaller groups should raise, not lower, per-person centrality—is plausible but untested against a size-matched human control. Within-treatment AI rank profile (RQ1) and the airtime–felt-value correlation are cleaner; the load-bearing causal leap is the between-condition human-human cost.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper reports a between-teams RCT (16 teams of 2 students + AI vs 17 teams of 3 students) on a chat-based moral-dilemma rescue task. Using Group Communication Analysis, post-task surveys, and lexical/temporal probes, it claims: (RQ1) the AI is consistently the most participatory and internally cohesive teammate while contributing the least new and least dense content; (RQ2) AI presence reduces human–human responsivity and social impact and lowers belonging and status, with AI airtime dominance correlating with lower felt value; (RQ3) this “social cost” is present at baseline rather than emerging within the session. The authors interpret the AI as a relational displacer and outline a broader design agenda.","tokens_in":13942,"tokens_out":1551,"duration_ms":33270,"significance":"If the human–human social-cost claim holds under cleaner controls, the paper makes a genuine contribution to human–AI teaming by shifting attention from trust-in-AI to effects on human–human relational fabric—an understudied channel. Strengths include a randomized design, clear RQs, within-team paired AI-vs-human contrasts for RQ1, effect sizes with a Bonferroni note on primary tests, a female-only sensitivity check, multi-method convergence (GCA + surveys + airtime–felt-value correlation), and an explicit research agenda. RQ1’s role profile is especially clean and useful for design. The construct of social cost is legible and falsifiable in follow-on work.","major_comments":[{"comment":"§3.1 and §3.5: The load-bearing RQ2 claim attributes lower human–human responsivity, social impact, belonging, and status to AI presence, but the design confounds condition with human group size (2 vs 3 students) and with GCA network composition (AI included as a third node vs three humans). Per-student scoring and within-pipeline z-standardization do not remove that structure: treatment humans have one human partner plus an AI node; controls have two human partners. The defense that smaller groups should raise per-person centrality is plausible but untested against a size-matched all-human dyad (or a 3-human + silent/observer control). Wheelan (2009) is cited yet not used as a design check. Without additional analyses or a matched control, causal attribution of the between-condition social cost specifically to AI presence remains underdetermined. Please either (a) add size-matched human","section":"§3.1, §3.5, §4.2"},{"comment":"§4.2.1 / §3.5: Overall responsivity and social impact are defined via uptake among teammates and are sensitive to who is in the discourse graph. Including the AI as a GCA node means treatment students’ relational scores partly reflect AI–human links, while the narrative claims reduced communication “with one another” (human–human). If human–human-only edge restriction was not the primary specification, report it; if it was, state the procedure and show that the between-condition gap survives when AI turns are excluded from the uptake graph. This is necessary for the relational-displacement interpretation.","section":"§4.2.1, §3.5"},{"comment":"§4.3: Interpreting four null temporal probes (all p > .12) as “positive evidence” that the cost is an immediate baseline property overclaims what nulls can support at N ≈ 33 teams. Low power, short sessions (~18 min), and length equalization can all yield non-emergence of a true dynamic. Reframe as failure to detect within-session emergence, report effect-size bounds or equivalence-style intervals where possible, and keep the baseline claim provisional pending longitudinal work already flagged in §5.2–5.3.","section":"§4.3, §5.2"},{"comment":"§4.2.2: Belonging (p = .011) does not survive the stated Bonferroni threshold (α = .0083) while responsivity, social impact, and status do. The text currently packages “four outcomes” with directional consistency; please separate corrected vs uncorrected results clearly in text and figures, and avoid implying uniform statistical support for the full survey bundle.","section":"§4.2.2"}],"minor_comments":[{"comment":"Figure 1 caption says values are “normalized within each dimension,” but the method of normalization (z within full sample vs within treatment only, etc.) is not stated in §3.4–3.5; add one sentence so the mid-range AI positions on responsivity/impact are interpretable.","section":"Figure 1, §3.4"},{"comment":"§3.1: Report total analytic N of students per condition after any exclusions, mean/SD conversation length, and AI turn share distribution; these are needed to contextualize participation and airtime dominance.","section":"§3.1"},{"comment":"§3.4: Belonging and status composites need item counts, response scales, and reliability (α/ω) in this sample; “adapted from Chung et al., 2020” is insufficient for reproduction.","section":"§3.4"},{"comment":"§3.5: GCA embedding model, turn-window size, and similarity threshold are described only as “established practice”; cite the exact parameter set or appendix them.","section":"§3.5"},{"comment":"Gender imbalance (63 F / 12 M; only three males in treatment) is acknowledged; state explicitly in §4 whether primary tables are full-sample or note the female-only sensitivity results in the main text, not only as a planned check.","section":"§3.4, §4"},{"comment":"Typos/style: “V oice” appears with a stray space in §5.3; arXiv IDs and “Advance online publication” entries are fine for a preprint but should be cleaned for journal production.","section":"§5.3, References"}],"recommendation":"major_revision","confidential_remarks":"The group-size/GCA-node confound is the decisive issue; if the authors can show human–human-only uptake gaps and/or add a size-matched control analysis from the broader TRAIL program, this could become a strong contribution. Novelty is real on the human–human channel but the manuscript leans heavily on same-lab companion citations (TRAIL, Choi et al., Park et al., Samadi et al.); that is acceptable if methods are self-contained. Fit for a serious cs.HC / teams venue is good after major revision. I would not accept on the current causal wording of RQ2."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core here is not another trust-in-AI finding. It is a six-dimension GCA portrait of the agent (most talkative and self-cohesive in every treatment team; lowest newness and density) plus converging signs that humans exchange less uptake with each other and report lower belonging/status when that agent is present. The airtime–felt-value correlation inside treatment teams is the cleanest human-side signal. Framing that pattern as human-human “social cost,” and checking whether it grows inside the session, is a real move for HAT design work that usually stops at human–AI trust and task output.\n\nWhat they did well: randomized teams, explicit RQs, paired within-team AI-vs-student tests for RQ1, effect sizes, a Bonferroni note on the primary between-condition tests, a female-only sensitivity check, and multi-method stacking (GCA, surveys, LIWC/LSM/RQA/stance probes). The AI profile is internally consistent and does not depend on the control arm. Citations sit on standard instruments (Dowell GCA, Chung belonging) rather than re-labeling noise.\n\nThe soft spot that matters is the one the stress-test flags. RQ2 compares 2 students+AI to 3 students. Per-student scoring and z-standardization do not erase unequal human partners or a three-node graph that includes the AI. Their defense—that smaller groups should raise centrality—is plausible and untested against a size-matched human dyad or a 3-human+AI add-on arm. Belonging sits at the edge of their own correction. “Immediate, not emergent” is mostly null temporal probes read as positive evidence; that is thin with N≈33 teams. Sample gender skew, single persona, text-only one-shot task, and no released prompts/code/data keep reproducibility modest.\n\nI would still send this to referees. The within-treatment role profile and the design question are worth the community’s time; a serious review should demand a size-matched control and materials, not a desk reject. For reading group: yes if people are building teammate agents; otherwise skim RQ1 and Figure 3. I would cite the AI GCA profile and the social-cost framing when discussing evaluation metrics beyond trust.","headline":"Clean within-team AI role profile and a real design question, but the between-condition “social cost” still rides on a group-size confound the paper cannot fully neutralize.","tokens_in":14672,"tokens_out":575,"would_cite":true,"duration_ms":17390,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"An AI teammate takes the most airtime and coheres with itself, yet its presence immediately lowers how human teammates respond to one another and how much they feel they belong.","keywords":["human-AI teaming","Group Communication Analysis","social cost","belonging","conversational agents","sociocognitive dynamics","team decision-making"],"falsifier":"Rerun the same task with human headcount matched—for example two humans versus two humans plus AI, or three humans versus three humans plus a controlled-airtime AI—and test whether human-human responsivity, social impact, belonging, and the airtime–felt-value correlation still appear only when the AI is present and talkative.","tokens_in":14472,"feed_emoji":"💬","tokens_out":958,"duration_ms":26751,"temperature":0.7,"pith_summary":"This paper argues that putting a conversational AI on a small decision-making team changes not only how people relate to the machine, but how they relate to each other. In a chat-based moral-dilemma task, the AI was the most talkative and self-referential member of every mixed team, while adding the least new and least dense content. Humans on those teams built on one another less, had less of their own contributions taken up, and reported lower belonging and status than people on all-human teams of three. The more the AI dominated airtime, the less valued students felt. Temporal probes tied to a mid-task moral reveal found no within-session emergence: the social cost was present from the start. The authors treat this as a baseline property of AI presence and open a design agenda aimed at protecting the human-human channel without discarding useful AI contribution.","feed_headline":"AI teammates cut human-to-human talk and belonging","feed_subtitle":"Students felt less valued as the AI took more chat airtime; the cost showed up from the first turns","key_machinery":"Group Communication Analysis (GCA): six sequential-semantic measures—participation, internal cohesion, overall responsivity, social impact, newness, and communication density—that locate each teammate’s sociocognitive role and quantify human-human uptake in multiparty chat.","core_discovery":"An AI teammate enacts a high-participation, high-internal-cohesion, low-newness, low-density role and imposes an immediate human-human social cost: students in AI teams show lower responsivity and social impact toward one another, report lower belonging and status, and feel less valued as AI airtime dominance rises. That cost is a baseline property of having the AI on the team rather than a dynamic that develops over the conversation.","pith_inferences":["If the cost is truly baseline and airtime-scaled, default high-verbosity agent settings in team products may systematically erode peer belonging even when users rate the agent as helpful.","Matching studies that hold human count fixed while varying only AI talkativeness would isolate displacement from group-size artifacts more cleanly than the present between-teams contrast.","The same GCA-plus-belonging package could benchmark facilitation modes (read-the-room versus lead-the-room) as a practical acceptance test before deployment in classrooms or workplaces.","Interaction with prior trust in AI is bidirectional in principle: high trust might free humans to talk to each other or instead cede still more relational attention to the agent."],"forward_implications":["Designers cannot judge an AI teammate only by human-AI trust or task output; they must also measure human-human uptake and belonging.","Calibrating AI airtime and relational behavior becomes a primary design lever if dominance scales felt devaluation.","Single-session text results motivate longitudinal and voice/embodied tests of whether the cost attenuates, compounds, or can be reversed.","Role designs that withdraw when human-human exchange is forming, or signal availability without absorbing attention, become concrete next experiments.","Social cost is separable from social loafing: substantive human content can hold while relational exchange falls."],"fun_headline_variants":["AI teammate dominates talk, cuts human responsivity and belonging","Humans talk less to each other when AI joins the team","AI airtime linked to students feeling less valued","AI teammate imposes immediate social cost on human pairs","Most talkative AI adds least new info, lowers human status"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That swapping one human for an AI (two students plus AI versus three students) does not itself drive the drop in human-human responsivity, belonging, and status beyond what per-student scoring can fix.","fun_headline_variants_meta":{"raw":{"variants":["AI teammate dominates talk, cuts human responsivity and belonging","Humans talk less to each other when AI joins the team","AI airtime linked to students feeling less valued","AI teammate imposes immediate social cost on human pairs","Most talkative AI adds least new info, lowers human status"]},"model":"grok-4.5","effort":"low","cost_usd":0.002504,"raw_usage":{"total_tokens":989,"prompt_tokens":803,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":25044000,"prompt_tokens_details":{"text_tokens":803,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":122,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":803,"tokens_out":64,"duration_ms":3188,"temperature":1.0,"reasoning_tokens":122,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T13:52:22.306696+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rerun the same task with human headcount matched—for example two humans versus two humans plus AI, or three humans versus three humans plus a controlled-airtime AI—and test whether human-human responsivity, social impact, belonging, and the airtime–felt-value correlation still appear only when the AI is present and talkative.","supporting_citations":[],"review_version":2}