Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

AI in Software Engineering: Perceived Roles and Their Impact on Adoption

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Developers who assign multiple roles to an AI coding tool report it as more useful and easier to use, and the paper treats this as a path to wider adoption.

desk verdict A useful role taxonomy for AI coding tools, but the paper's central claim that 'diverse conceptualizations enhance adoption' is not supported by the reported measure, which counts roles rather than diversity. read the letter →

arxiv 2504.20329 v1 pith:35OIU2PM submitted 2025-04-29 cs.SE cs.HC

classification cs.SEcs.HC
keywords AI4SEmentalmodelsroleattributiontechnologyacceptancemodelperceivedusefulnesseaseofusedevelopersurveyfactoranalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks how developers mentally frame AI-powered development tools and whether that framing predicts adoption. Drawing on 38 interviews and a survey of 102 developers, it identifies two broad mental models—AI as an inanimate tool and AI as a human-like teammate—and shows statistically that the roles developers assign group into Support Roles and Expert Roles. Its central finding is that the more roles a developer assigns to AI, the higher their Perceived Usefulness and Perceived Ease of Use, which the authors read as evidence that diverse conceptualizations support adoption. If true, this gives concrete design and onboarding levers for AI coding tools: help developers see the tool in multiple roles rather than a single narrow one.

What carries the argument

The central object is the role attribution itself: the set of roles a developer says AI plays, measured by a thirteen-option survey derived from a yearly developer ecosystem report. The argument is carried by two statistical structures built on those attributions—a two-factor model distinguishing Expert Roles from Support Roles, and Pearson correlations of role counts with Perceived Usefulness and Perceived Ease of Use from a Revised TAM questionnaire. The factor model and the correlation table are what turn qualitative talk about 'assistant' or 'colleague' into a claim about adoption.

What would settle it

A replication using a representative sample of developers recruited outside an opt-in user-study panel, allowing open-ended role descriptions instead of a fixed list, would falsify the claim if the number of attributed roles showed no positive correlation with perceived usefulness and perceived ease of use.

Watch

Extended reading notes

Core claim

The paper's central claim is that role attribution is not a passive byproduct of using AI tools but a measurable factor in technology acceptance. In the survey data, the total number of roles assigned correlates with both TAM scales ($r = 0.59$ with perceived usefulness, $r = 0.56$ with perceived ease of use, both $p < 0.001$), and both factor-derived role dimensions correlate positively with acceptance. Factor analysis with varimax rotation and a $>0.4$ loading threshold splits the thirteen offered roles into Expert Roles (advisor, reviewer, problem solver) and Support Roles (assistant, reference guide, tool). The authors interpret this as showing that developers who conceptualize AI along multiple dimensions—both helper and expert, both tool and teammate—find it more useful and easier to integrate, and they argue this is consistent with the qualitative divide between tool-minded and teammate-minded developers.

Load-bearing premise

The load-bearing premise is that the 102 people on the tool maker's opt-in study list, and the thirteen role labels the survey offered them, represent the broader population of developers and the roles developers actually attribute to AI.

Editorial extensions

If this is right

  • If role count predicts perceived usefulness and ease of use, designers can nudge adoption by helping developers see a tool in more than one role, for instance as both assistant and reviewer.
  • Because Support and Expert role attributions load on separate factors, tools that support both kinds of framing may cover a wider range of developer expectations than tools aimed at a single role.
  • The qualitative pattern—tool-minded developers enforcing stricter technical standards while teammate-minded developers tolerate imperfections—implies that a single onboarding message will not fit all users.
  • The positive correlation between role attribution and the number of AI tools tried suggests that exposure and conceptualization may reinforce each other in an adoption cycle.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The observed correlation may partly reflect a general AI enthusiasm factor: developers who like AI might both assign more roles and rate the tool more favorably, a possibility the correlational design cannot separate.
  • A testable extension would be to experimentally prime a single role, such as 'junior colleague' versus 'tool', and measure whether actual task performance and sustained usage shift, rather than only self-reported perceptions.
  • The two mental models may correspond to different trust-calibration strategies: a tool framing may invite verification and reduce over-reliance, while a teammate framing may increase tolerance but risk over-trust—a trade-off the paper leaves implicit.
  • The interview pattern in which novices described AI as a teacher and experienced developers as a junior engineer suggests a developmental trajectory for role attribution that longitudinal studies could track.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper investigates how developers conceptualize AI-powered development tools and whether these role attributions relate to technology acceptance. It combines a secondary qualitative analysis of 38 interviews with a new survey of 102 participants, identifies two role dimensions via factor analysis (Support Roles and Expert Roles), and reports positive correlations between the number of assigned roles, Perceived Usefulness, and Perceived Ease of Use. The authors interpret these correlations as evidence that diverse conceptualizations of AI enhance adoption and propose adaptive design and onboarding strategies for AI4SE tools.

Significance. If the central claim holds, the paper would provide actionable guidance for designing AI4SE tools that accommodate different user mental models. The study has clear strengths: the qualitative grounding in interview data, the transparent reporting of the survey instrument via a DOI, and the inclusion of a correlation table with significance levels. The two-factor structure of roles is plausible and consistent with prior work on mental models of AI. However, the load-bearing interpretation that 'diverse conceptualizations enhance AI adoption' is not directly supported by the measured variable (total number of roles), and the causal wording exceeds what cross-sectional correlational data can establish.

major comments (3)
  1. [Section 4.2, Table 1] The central claim equates the total number of assigned AI roles with diversity of mental models, but Table 1 only reports correlations with role count. A participant can endorse many Support Roles (assistant, tool, reference guide, content generator) and zero Expert Roles, yielding a high role count but a single-dimensional, homogeneous mental model. Conversely, endorsing one Support and one Expert role yields a low count but a genuinely diverse conceptualization. The factor analysis shows the Support and Expert factors are negatively correlated (-0.15, not significant), so within-cluster endorsement is entirely plausible. The paper never computes a diversity index, cross-category endorsement, or an interaction between the two factors. Therefore the observed correlation could be driven by general AI enthusiasm, acquiescence, or depth of engagement rather than by the proposed diversity mechanism. This is a construct-validity gap at the center of the paper and needs to be addressed either by re-analyzing the data with a proper diversity measure or by reframing the conclusion to 'number of roles' rather than 'diverse conceptualizations.'
  2. [Abstract and Section 5] The statements 'diverse conceptualizations enhance AI adoption' and 'Mental Models of AI directly influence technology adoption decisions' imply a causal direction, but the study is a cross-sectional survey. The correlation between role count and PU/PEU could reflect reverse causality (users who find a tool useful may be motivated to explore and assign more roles to it) or a third variable such as general engagement with AI tools. The paper should soften the causal language and, at minimum, control for the number of AI tools tried and coding experience in a regression or partial correlation analysis, since these variables are already measured and reported in Table 1.
  3. [Section 3] The survey sample is a convenience sample of 102 participants recruited from JetBrains' curated list of people who had previously consented to user studies, and the role options were derived from JetBrains' Developer Ecosystem Report. This dual dependence on JetBrains-affiliated channels may limit the representativeness of both the role distribution and the correlations with acceptance. The authors should explicitly discuss this limitation and temper the generalizability claims, or provide evidence that the sample is diverse in terms of tool usage and professional background beyond the reported experience levels.
minor comments (5)
  1. [Table 1] The table header contains a typo: 'AT tools tried' should be 'AI tools tried'.
  2. [Section 4.2] The text mentions 'teacher, mentor, senior colleague, or junior colleague' as roles, but the survey options listed in Section 3 include 'teacher' but not 'mentor' (the closest option is 'senior colleague' or 'companion'). Please align the description with the actual survey options.
  3. [Section 4.2] The factor analysis section reports factor loadings but does not report eigenvalues, the proportion of variance explained, or the scree plot itself. Adding these would strengthen the justification for choosing two factors.
  4. [Section 4.2] The sentence '24+ with 16 and more years of experience' is awkwardly phrased; '24 participants with 16 or more years' would be clearer.
  5. [Section 3] The reference to the survey DOI is good, but the paper should also state whether the anonymized dataset and analysis scripts are available, since the manuscript says 'available upon request' in Section 4, which is less transparent than the survey instrument.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central PU/PEU correlations come from new survey data; the self-cited interview re-analysis is transparent and not used as a fitted input.

full rationale

The paper's central quantitative claim is the correlation between the number of AI roles assigned and PU/PEU (Section 4.2, Table 1). These correlations are computed from the 102-participant survey; PU and PEU are measured with the Revised TAM questionnaire and are not derived from role counts or factor loadings. The factor analysis is used to summarize role structure, but the factor scores are independent variables correlated with PU/PEU, not predicted outcomes. The qualitative Mental Models are explicitly labeled as a re-analysis of the authors' prior study [15], and the current paper does not present that reuse as an external, independent proof; it is transparent secondary analysis. The survey's role options come from the JetBrains Developer Ecosystem Report [9], which is an input to instrument design rather than a result that predetermines the correlation. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported, and no known result is merely relabeled. The abstract's phrase 'diverse conceptualizations' may overstate what role count measures, but that is a construct-validity concern about interpretation, not a circular derivation. Therefore no specific circular step could be quoted and reduced by construction.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim rests on a convenience sample, a researcher-chosen factor count and loading threshold, and standard but unvalidated-in-context TAM scales. No new entities or parameters are postulated beyond the factor analytic choices.

free parameters (2)
  • Number of factors extracted = 2
    Chosen by visual inspection of the scree plot (Section 4.2). This determines the Support vs Expert role grouping and is a researcher judgment, not a parameter from a prior model.
  • Factor loading significance threshold = 0.4
    Loadings above 0.4 were interpreted as significant, following Rogers (2022). This threshold controls which roles load on which factor.
assumptions (4)
  • domain assumption The 'define AI' question in the original interviews is a valid and reliable probe of each participant's mental model of AI-powered development tools.
    The qualitative 80/20 split (Section 4.1) is a secondary analysis of a single interview question from the authors' prior study [15]; no coding scheme, inter-rater reliability, or validation is reported.
  • domain assumption The 102 survey participants are representative of the broader developer population.
    Recruitment was from JetBrains' curated list of user-study volunteers (Section 3), which likely over-represents JetBrains product users and people with a prior relationship to the company, limiting generalizability.
  • domain assumption The Revised TAM Questionnaire is reliable and valid for AI-powered development tools in this sample.
    No Cronbach's alpha or validity checks are reported; the near-perfect correlation between PU and PEU (r = 0.94) suggests possible common-method bias or scale redundancy.
  • domain assumption Pearson correlation is appropriate for the count of roles and Likert-scale variables.
    The total number of roles is a count and individual roles are binary selections; treating these as interval data for Pearson correlation is an unstated assumption that can affect significance levels.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI in Software Engineering: Perceived Roles and Their Impact on Adoption." pith.science (2026). https://pith.science/paper/35OIU2PM

@misc{pith2026250420329,
  author       = {Pith},
  title        = {Pith review of: AI in Software Engineering: Perceived Roles and Their Impact on Adoption},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/35OIU2PM}},
  note         = {Machine review of arXiv:2504.20329}
}
read the original abstract

This paper investigates how developers conceptualize AI-powered Development Tools and how these role attributions influence technology acceptance. Through qualitative analysis of 38 interviews and a quantitative survey with 102 participants, we identify two primary Mental Models: AI as an inanimate tool and AI as a human-like teammate. Factor analysis further groups AI roles into Support Roles (e.g., assistant, reference guide) and Expert Roles (e.g., advisor, problem solver). We find that assigning multiple roles to AI correlates positively with Perceived Usefulness and Perceived Ease of Use, indicating that diverse conceptualizations enhance AI adoption. These insights suggest that AI4SE tools should accommodate varying user expectations through adaptive design strategies that align with different Mental Models.

Figures

Figures reproduced from arXiv: 2504.20329 by the authors.

Figure 1
Figure 1. Hierarchical Clustering of the roles that developers assign to AI [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimizing Token Consumption in LLMs: A Nano Surge Approach for Code Reasoning Efficiency

    cs.SE 2025-04 conditional novelty 3.0 of 10

    Refactoring smelly Java code and adding context, role, or token-limit prompts cut LLM chain-of-thought token use by roughly 15-50% in this study, but the 'no quality loss' claim rests only on shallow similarity metrics.

Reference graph

Works this paper leans on

21 extracted references · 12 canonical work pages · cited by 1 Pith paper

  1. [1]

    Astobiza

    Aníbal M. Astobiza. 2024. Do people believe that machines have minds and free will? Empirical evidence on mind perception and autonomy in machines. AI and Ethics 4, 4 (2024), 1175–1183

  2. [2]

    Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S Lasecki, Daniel S Weld, and Eric Horvitz. 2019. Beyond accuracy: The role of mental models in human-AI team performance. In Proceedings of the AAAI conference on human computation and crowdsourcing, Vol. 7. 2–11

  3. [3]

    Jeremy Biggs and Nitin Madnani. 2024. factor-analyzer: A Python package for factor analysis. https://github.com/EducationalTestingService/factor_analyzer. Version 0.5.1, accessed March 5, 2025

  4. [4]

    Rudrajit Choudhuri, Bianca Trinkenreich, Rahul Pandita, Eirini Kalliamvakou, Igor Steinmacher, Marco Gerosa, Christopher Sanchez, and Anita Sarma. 2024. What Guides Our Choices? Modeling Developers’ Trust and Behavioral Intentions Towards GenAI. arXiv preprint arXiv:2409.04099 (2024)

  5. [5]

    Cursor. 2025. Cursor AI-powered code editor . https://www.cursor.com/ Accessed: March 5, 2025

  6. [6]

    Daniel C Dennett. 1989. The intentional stance. MIT press

  7. [7]

    GitHub. 2025. GitHub Copilot. https://github.com/features/copilot Accessed: March 5, 2025

  8. [8]

    ICC/ESOMAR. [n. d.]. International Code on Market, Opinion and So- cial Research and Data Analytics. https://esomar.org/uploads/attachments/ ckqtawvjq00uukdtrhst5sk9u-iccesomar-international-code-english.pdf. Ac- cessed: October 2024

Show all 21 references
  1. [9]

    JetBrains. 2024. The State of Developer Ecosystem 2024. https://www.jetbrains. com/lp/devecosystem-2024/ Accessed: February 13, 2025

  2. [10]

    (Jim) Lewis

    James R. (Jim) Lewis. 2018. Comparison of Four TAM Item Formats: Effect of Response Option Labels and Order. UXPA Journal 14, 4 (2018), 224–236. https://uxpajournal.org/tam-formats-effect-response-labels-order/

  3. [11]

    Ze Shi Li, Nowshin Nawar Arony, Ahmed Musa Awon, Daniela Damian, and Bowen Xu. 2024. AI tool use and adoption in software development by individuals and organizations: a grounded theory study. arXiv preprint arXiv:2406.17325 (2024)

  4. [12]

    Pat Pataranutaporn, Ruby Liu, Ed Finn, and Pattie Maes. 2023. Influencing human–AI interaction by priming beliefs about AI can increase perceived trust- worthiness, empathy and effectiveness. Nature Machine Intelligence 5, 10 (Oct. 2023), 1076–1086. doi:10.1038/s42256-023-0072...

  5. [13]

    Pablo Rogers. 2022. Best practices for your exploratory factor analysis: A factor tutorial. Revista de Administração Contemporânea 26, 06 (2022), e210085

  6. [14]

    Daniel Russo. 2024. Navigating the complexity of generative ai adoption in software engineering. ACM Transactions on Software Engineering and Methodology 33, 5 (2024), 1–50

  7. [15]

    Sergeyuk, E

    A. Sergeyuk, E. Koshchenko, I. Zakharov, T. Bryksin, and M. Izadi. 2024. The Design Space of in-IDE Human-AI Experience. arXiv preprint arXiv:2410.08676 (2024). https://arxiv.org/abs/2410.08676

  8. [16]

    Nancy Staggers and Anthony F. Norcio. 1993. Mental models: concepts for human-computer interaction research. International Journal of Man-machine studies 38, 4 (1993), 587–605

  9. [17]

    Tabnine. 2025. Tabnine AI Code Completion. https://www.tabnine.com/ Accessed: March 5, 2025

  10. [18]

    Viswanath Venkatesh and Fred D Davis. 2000. A theoretical extension of the technology acceptance model: Four longitudinal field studies.Management science 46, 2 (2000), 186–204

  11. [19]

    Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J

    Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jar- rod Millman, Nikolay Mayorov, Andrew R. J. Nelson,...

  12. [20]

    Qiaosi Wang and Ashok K. Goel. 2022. Mutual theory of mind for human-AI communication. arXiv preprint arXiv:2210.03842 (2022)

  13. [21]

    Shao Zhang, Xihuai Wang, Wenhao Zhang, Yongshan Chen, Landi Gao, Dakuo Wang, Weinan Zhang, Xinbing Wang, and Ying Wen. 2024. Mutual theory of mind in human-ai collaboration: An empirical study with llm-driven ai agents in a real-time shared workspace task. arXiv preprint arXiv...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.