{"id":"ecf2813d-0faa-4883-b112-f10a7db48546","arxiv_id":"2509.10956","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A two-year interview study found generative AI boosted individual speed but did not fix team coordination, while shifting team culture toward efficiency and transparent AI use.","lead":"A two-year interview study of 15 software professionals found that AI did not solve core teamwork problems such as accountability and communication, but it changed team culture: efficiency became expected and transparent AI use became a professional norm.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Protocol reminder may inflate culture-shift findings; paper's mitigation logic is inverted","rationale":"Reader correctly identified the retrospective-report concern as the weak point. I agree and sharpen it: the paper's design actually makes the threat worse than 'could steer'; Section 3.3's claim that reminding participants of 2023 statements mitigates prompt artifacts is inconsistent with basic demand-characteristic reasoning. Reminder-based elicitation invites participants to construct a 'then vs. now' story, which is exactly the trajectory the paper reports. This concern is load-bearing because the culture-shift claim is the paper's most novel contribution; without it, the paper mostly confirms known results about individual productivity gains. The concrete test—coding occurrence position relative to reminders—is feasible with existing data and would settle whether the themes arise independently. The verdict stays CONDITIONAL because the paper can address this with reanalysis and transparent reporting; it does not need rejection.","tokens_in":28337,"tokens_out":5331,"duration_ms":61704,"concrete_test":"Re-analyze the 2025 interview transcripts and code, for each of the three central culture-shift themes (efficiency norms, transparency/responsible-use professionalism, AI acceptance), the first occurrence and the immediately preceding interviewer question. Determine whether the first occurrence comes in response to a question that references the participant's 2023 statements or before any such reference. If the themes appear in answers to generic current-practice questions (e.g., 'How do you use AI now?' or 'How has your team changed?') before any reminder, the concern is mitigated. If they emerge only after the interviewer explicitly brings up 2023 statements, the culture-shift conclusion should be weakened or re-presented as a co-constructed account. Report a table per participant and theme with the prompt type.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The single most load-bearing concern is that the central culture-shift finding may be an artifact of the Wave 2 interview protocol, and the paper's stated mitigation is internally misconceived. Section 3.3 says: 'To mitigate the risk that this contrast was simply a prompt artifact, we reminded participants of their 2023 statements before inviting comparison.' Reminding participants of their past hopeful statements immediately before asking them to talk about the present is a textbook contrast prime; it does not mitigate prompt artifacts, it amplifies them. The finding that AI had not fixed teamwork and instead shifted culture—with participants emphasizing unfulfilled hopes and new efficiency norms—could therefore be an expected response to the reminder, not a spontaneously offered longitudinal assessment. The paper offers no behavioral trace data (e.g., Slack logs, project artifacts, meeting summaries) to independently confirm either the persistence of the collaboration problems or the emergence of new cultural norms. The Limitation section (Section 7) does acknowledge that narratives 'might reflect... interpretive frames that we could have co-constructed with them during the very process of interviewing,' but it does not fix the earlier mischaracterization of the reminder as a mitigation. If this concern lands, the most novel part of the claim—that AI shifted collaborative culture—is not supported by the evidence as reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a two-wave qualitative longitudinal interview study (2023 and 2025) of 15 members (10 re-interviewed) of a distributed, project-based software development organization. In early 2023, participants imagined AI as an intelligent coordinator that could improve accountability and communication. By 2025, participants described AI as mainly an individual productivity tool, with core teamwork problems persisting, while collaborative culture had shifted: efficiency became expected, transparent and responsible AI use became markers of professionalism, and AI was normalized in teamwork. The paper frames this as evidence that AI has not fixed teamwork but has reshaped collaborative culture, and offers design implications for future teamwork AI.","tokens_in":28620,"tokens_out":3068,"duration_ms":37727,"significance":"If the central claim holds, this is a valuable longitudinal empirical contribution to CSCW/HCI and the broader AI-and-work literature: it challenges both techno-optimist and techno-pessimist narratives by showing that AI's team-level effect can be cultural normalization rather than coordination repair. The paper's strengths include a rare two-wave panel design at a strategically relevant inflection point (2023–2025), explicit use of sociotechnical imaginaries and domestication theory, and a candid limitations section that acknowledges the absence of behavioral data and the possibility of co-constructed narratives. The study is also helpfully explicit about its scope: 15 self-selected participants from one organization. The authors do not overclaim statistical generalizability. However, the novelty of the culture-shift claim rests heavily on the design of the Wave 2 protocol, and that design has a critical flaw discussed below.","major_comments":[{"comment":"The stated mitigation for prompt artifacts is logically inverted. The paper says: 'To mitigate the risk that this contrast was simply a prompt artifact, we reminded participants of their 2023 statements before inviting comparison.' Reminding participants of their earlier optimistic statements immediately before asking them to evaluate the present is a textbook contrast prime; it can inflate perceived change and disappointment, not reduce it. The limitation in Section 7 acknowledges that narratives 'might reflect... interpretive frames that we could have co-constructed,' but that does not repair the mischaracterization of the reminder as a mitigation. This is load-bearing because the central culture-shift claim depends on the before/after contrast. Please reframe the procedure as a limitation, and either add triangulating behavioral evidence (e.g., logs, outputs) or soften the causal clai","section":"Section 3.3"},{"comment":"The Wave 2 protocol explicitly introduced topics that were not part of the Wave 1 protocol: participants were asked about 'any shifts in team culture, such as norms around efficiency, or accountability' (Section 3.3). Phase 1 (Section 3.2) asked about frictions, workarounds, and forward-looking imaginaries—not about existing cultural norms. The 2025 themes in Section 5.3 (efficiency as a norm, transparency as professionalism, AI as expected) may therefore be an artifact of newly introduced questions rather than independently observed longitudinal change. The reminder of 2023 statements does not remedy this asymmetry. To support the longitudinal claim, the paper should either provide evidence that these norm-related themes were absent from comparable 2023 data (e.g., a systematic comparison of both transcripts for each participant), or reframe the finding as an emergent, retrospectively c","section":"Section 3.3 and Section 5.3"}],"minor_comments":[{"comment":"Typo: 'perfceptions' should be 'perceptions'.","section":"Section 3.2"},{"comment":"Typo: 'commutation' should be 'communication' in 'AI for enhancing commutation and maintaining team relationships'.","section":"Section 4.2.3"},{"comment":"Typo: 'doption' should be 'adoption' in 'Most AI doption research'.","section":"Section 2.1"},{"comment":"The sentence 'To short, our follow-up interviews in section 5 confirmed...' should read 'In short, our follow-up interviews in Section 5 confirmed...'.","section":"Section 5.3.3"},{"comment":"Typo: 'highligting' should be 'highlighting' in design implication 1.","section":"Section 6.4"},{"comment":"Typo: 'unobstrusive' should be 'unobtrusive'.","section":"Section 7"},{"comment":"The reference format line contains 'In.ACM' with a stray period; please fix to 'In Proceedings of the ACM Conference...'.","section":"ACM Reference Format"}],"recommendation":"major_revision","confidential_remarks":"The core empirical contribution is timely and the qualitative data are presented with appropriate caution overall. However, the central culture-shift claim rests on a methodological design whose stated mitigation is logically flawed. This is fixable by reframing the result as a retrospective, participant-constructed account and by strengthening the limitations discussion, but as written the paper overstates the evidentiary basis for 'AI shifted collaborative culture.' The paper would also benefit from reporting more of the coding process (e.g., number of coders, discussion of divergent themes) even if reflexive thematic analysis does not require inter-rater reliability. This is not a reject; the longitudinal dataset is valuable and the authors have the raw material to recalibrate the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a rare and useful longitudinal qualitative dataset, and the culture-shift claim is plausible. But the Wave 2 protocol has a real flaw that the paper mis-describes as a mitigation, and the central before/after comparison is the part most affected.\n\nThe new thing: a two-wave interview study spanning the ChatGPT inflection point (2023-2025) in a distributed software organization, with 10 of 15 participants returning. That alone is scarcer than it should be. The finding that AI did not fix performance accountability or communication, but did normalize efficiency and make transparency a professional norm, is a genuinely useful corrective to the hype. The paper connects this to sociotechnical imaginaries and domestication theory without overreaching. The limitations section is candid, and there is no modeling or fitting to worry about.\n\nThe soft spot: Section 3.3 says the authors reminded participants of their 2023 statements \"to mitigate the risk that this contrast was simply a prompt artifact.\" In my reading, that's not a mitigation; it's a contrast prime. If you remind someone of their earlier hopeful projections and then ask them to compare, you increase the chance they'll frame the present as a disappointment. The stress-test note is right about that. The paper's own limitation section does acknowledge that narratives could be co-constructed in the interview, which suggests the authors know the risk—but the earlier sentence mischaracterizes how the design works.\n\nHow much does this matter? The 'persistent challenges' half of the story is directly vulnerable to the priming. The culture-shift half (efficiency norms, responsible-use professionalism, AI as taken-for-granted) is based largely on participants' descriptions of current practice, so it is less compromised. I don't think the finding is an artifact, but the before/after contrast should be treated with real caution. A revision should reframe that reminder as a limitation and ideally analyze the 2025 interviews without reference to the 2023 statements before doing the paired comparison.\n\nThe other limitations—15 people, one self-selected org, self-report without logs—are real but standard for qualitative longitudinal work. They don't sink the paper.\n\nBottom line: send it to peer review. It deserves referee time, and a good reviewer can push for the protocol fix. I'd probably cite it for the culture-shift observation, and I'd consider bringing it to a reading group to discuss the method, not just the findings.","headline":"A rare two-wave qualitative study that contributes a plausible culture-shift finding, but the Wave 2 reminder protocol is a contrast prime in disguise, not the mitigation the paper claims.","tokens_in":29064,"tokens_out":2876,"would_cite":true,"duration_ms":36556,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Over two years, AI became a personal booster, not a team coordinator, leaving collaboration problems unresolved while shifting team culture around efficiency and transparency.","keywords":["AI and teamwork","longitudinal interviews","collaborative culture","sociotechnical imaginaries","domestication theory","generative AI at work","distributed software teams","human-AI collaboration"],"falsifier":"Gather digital-trace data (chat logs, issue-tracker timelines, commit histories, meeting records) from a comparable distributed software organization across 2023–2025 and measure whether coordination failures—silent disengagement, late handoffs, duplicated work, communication fragmentation—actually remained constant or declined while AI adoption grew. If coordination failures show improvement, the claim that AI left core teamwork problems unfixed would be refuted.","tokens_in":28274,"feed_emoji":"🤖","tokens_out":4527,"duration_ms":50046,"temperature":0.7,"pith_summary":"This paper claims that AI's impact on teamwork in a distributed, project-based software organization was twofold: it largely failed to fix the collaboration problems participants hoped it would solve, and it reshaped the culture of teamwork in ways that persisted. The authors interviewed the same people in early 2023, when ChatGPT had just appeared, and again in 2025, after generative AI became common. In 2023, participants envisioned AI as an intelligent coordinator that would track progress, flag disengagement, and ease communication frictions. By 2025, AI was used mostly as an individual assistant for coding, writing, and documentation; accountability and communication breakdowns remained. What changed was cultural: efficiency became a norm, transparent and responsible AI use became a mark of professionalism, and AI became a taken-for-granted part of collaboration.","feed_headline":"AI didn't fix teamwork, but it shifted team culture","feed_subtitle":"Two-year interviews show AI stayed a personal booster while norms of efficiency and transparency took hold.","key_machinery":"The central machinery is the longitudinal qualitative interview design paired across two waves: each participant's 2023 and 2025 transcripts are analyzed as linked pairs to trace continuity and change in expectations, practices, and norms. The analysis is organized by two conceptual lenses—sociotechnical imaginaries (collectively held visions of desirable technological futures) and domestication theory (how technologies are appropriated, incorporated, and normalized into daily routines). These lenses let the authors treat the difference between imagined and actual AI use not as noise but as a meaningful trajectory: ambitious group-level hopes were domesticated into individual productivity ha","core_discovery":"The paper reports a two-wave longitudinal interview study with 15 members of a remote-first, project-based software development organization in early 2023 and 10 of the same people in 2025. In 2023, participants dreamed of AI as an ambient coordinator—flagging underperformance, tracking milestones, sensing communication breakdowns, and mediating relational friction. By 2025, they were using AI mainly to accelerate individual tasks like coding, writing, and documentation, while the collaboration problems of accountability and communication persisted. The authors argue that AI's main team-level effect was cultural rather than functional: efficiency became an expected norm, transparency and res","pith_inferences":["An implication the paper leaves implicit: a 'productivity trap' may form at team level—if AI-fueled speed becomes the baseline, those who cannot or will not use AI could be penalized, creating a new form of inequality inside teams.","A testable extension the authors do not explore: agentic AI systems that proactively monitor, summarize, and nudge team activity might reactivate the 2023 coordinator imaginaries and actually reduce coordination failures, or they might reproduce the same individual-speed-at-team-cost pattern at higher velocity.","Methodologically, the paper's reminder-based interview design is itself an intervention; a future study could vary whether and how prior statements are invoked to estimate how much of the reported 'hope-to-disappointment' trajectory is co-constructed by the recall process."],"forward_implications":["Individual productivity gains from AI do not automatically translate into better teamwork; without deliberate design for group awareness, the same coordination failures persist.","Team culture is a site where AI exerts measurable influence even when tools are used individually: efficiency expectations rise, transparency becomes a professional virtue, and AI use becomes normalized.","Future workplace AI should be designed for proactive sensemaking—flagging slow progress, shifts in tone, or disengagement before crisis—rather than reactive note-taking or summarization.","Teams need mechanisms to make cultural shifts visible and negotiable, so that implicit new norms around speed and AI use do not silently create misalignment or mistrust.","The 2023 imaginaries of AI as coordinator and mediator remain design opportunities; the study suggests they were displaced but not invalidated by current tool affordances."],"fun_headline_variants":["AI won't save teamwork, but it reshaped team culture","AI didn't fix collaboration, but it changed team norms","Teamwork problems persist as AI boosts individual work","AI's real team impact: new norms of efficiency and transparency","Two years with AI: teamwork unchanged, culture shifted"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The study's central trajectory—hopes for a collaborative coordinator giving way to individual productivity plus cultural normalization—rests entirely on what participants said retrospectively in interviews, since no behavioral or observational data were collected to verify actual changes in teamwork.","fun_headline_variants_meta":{"raw":{"variants":["AI won't save teamwork, but it reshaped team culture","AI didn't fix collaboration, but it changed team norms","Teamwork problems persist as AI boosts individual work","AI's real team impact: new norms of efficiency and transparency","Two years with AI: teamwork unchanged, culture shifted"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000873,"raw_usage":{"total_tokens":3585,"prompt_tokens":680,"completion_tokens":2905,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":424,"completion_tokens_details":{"reasoning_tokens":2835}},"tokens_in":424,"tokens_out":2905,"duration_ms":23890,"temperature":1.0,"reasoning_tokens":2835,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T17:18:10.447107+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Gather digital-trace data (chat logs, issue-tracker timelines, commit histories, meeting records) from a comparable distributed software organization across 2023–2025 and measure whether coordination failures—silent disengagement, late handoffs, duplicated work, communication fragmentation—actually remained constant or declined while AI adoption grew. If coordination failures show improvement, the claim that AI left core teamwork problems unfixed would be refuted.","supporting_citations":[],"review_version":1}