{"id":"52010b8c-fa11-44a1-b50f-18a0207f5c02","arxiv_id":"2504.20329","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Developers who attribute a greater number of roles to AI coding assistants, from tool to expert, report higher perceived usefulness and ease of use, but the evidence is correlational.","lead":"This paper reports on how developers perceive AI coding tools, identifying two main mental models: AI as an inanimate tool and AI as a human-like teammate. A survey of 102 people found that developers who assign more roles to AI also rate it as more useful and easier to use, hinting that flexible role framing may support adoption.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper equates 'number of assigned roles' with 'diverse conceptualizations,' but Table 1 only tests role count; a high count can come from endorsing many Support roles alone, so the diversity mechanism is untested.","rationale":"The reader identified sampling representativeness as the weakest assumption. That is a valid external-validity concern, but the more immediate threat is internal construct validity: the independent variable actually tested (total role count) does not correspond to the construct named in the central claim (diversity of conceptualizations). This issue would persist even with a perfectly representative sample, and it directly affects the design recommendations (adaptive modes, onboarding framing). The qualitative 80/20 split and the negative correlation between the two factor scores suggest that many participants cluster within one factor, so high counts can reflect within-cluster breadth rather than cross-cluster diversity. The paper should either reanalyze with a diversity measure or soften the interpretation to 'assigning more roles is associated with higher acceptance,' dropping the diversity mechanism. Because the data already exist, this is a feasible revision; the concern reinforces the reader's CONDITIONAL verdict rather than changing it, so verdict_should_be is UNCHANGED. I disagree with the reader's choice of weakest assumption because the construct-validity gap is more load-bearing for the paper's central interpretive claim.","tokens_in":6921,"tokens_out":5502,"duration_ms":57706,"concrete_test":"Reanalyze the item-level survey data (the authors state it is available on request) to construct a diversity measure, e.g., an indicator for endorsing at least one Expert Role and at least one Support Role, or an entropy score over the two factor clusters. Fit a regression of PU (and PEU) on total role count, the diversity indicator, and their interaction. If the diversity indicator or interaction is not significant after controlling for count, the central claim's 'diverse conceptualizations' mechanism is unsupported; if diversity adds predictive power, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 and Table 1 show that the total number of AI roles correlates with PU (r=.59) and PEU (r=.56), and the abstract interprets this as 'diverse conceptualizations enhance AI adoption.' The load-bearing assumption is that role count measures conceptual diversity. It does not: a participant can select four Support roles (assistant, tool, reference guide, content generator) and zero Expert roles, yielding a high count but a narrow, single Mental Model; another can select one Support and one Expert role, yielding a low count but genuinely diverse conceptualizations. The factor analysis finds Support and Expert role factors that are negatively correlated (-0.15, not significant), so endorsing many roles within a single cluster is plausible and perhaps common. The paper never computes a diversity index, cross-category endorsement, or an interaction between the two factors. Consequently, the observed correlation could be driven by general AI enthusiasm, acquiescence, or depth of engagement rather than by the diversity of Mental Models that the design recommendations depend on. The causal wording ('enhance') also exceeds the cross-sectional correlation, but the more specific problem is that the independent variable does not isolate the proposed mechanism. This is a construct-validity gap at the center of the paper.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how developers conceptualize AI-powered development tools and whether these role attributions relate to technology acceptance. It combines a secondary qualitative analysis of 38 interviews with a new survey of 102 participants, identifies two role dimensions via factor analysis (Support Roles and Expert Roles), and reports positive correlations between the number of assigned roles, Perceived Usefulness, and Perceived Ease of Use. The authors interpret these correlations as evidence that diverse conceptualizations of AI enhance adoption and propose adaptive design and onboarding strategies for AI4SE tools.","tokens_in":7071,"tokens_out":2306,"duration_ms":24045,"significance":"If the central claim holds, the paper would provide actionable guidance for designing AI4SE tools that accommodate different user mental models. The study has clear strengths: the qualitative grounding in interview data, the transparent reporting of the survey instrument via a DOI, and the inclusion of a correlation table with significance levels. The two-factor structure of roles is plausible and consistent with prior work on mental models of AI. However, the load-bearing interpretation that 'diverse conceptualizations enhance AI adoption' is not directly supported by the measured variable (total number of roles), and the causal wording exceeds what cross-sectional correlational data can establish.","major_comments":[{"comment":"The central claim equates the total number of assigned AI roles with diversity of mental models, but Table 1 only reports correlations with role count. A participant can endorse many Support Roles (assistant, tool, reference guide, content generator) and zero Expert Roles, yielding a high role count but a single-dimensional, homogeneous mental model. Conversely, endorsing one Support and one Expert role yields a low count but a genuinely diverse conceptualization. The factor analysis shows the Support and Expert factors are negatively correlated (-0.15, not significant), so within-cluster endorsement is entirely plausible. The paper never computes a diversity index, cross-category endorsement, or an interaction between the two factors. Therefore the observed correlation could be driven by general AI enthusiasm, acquiescence, or depth of engagement rather than by the proposed diversity mechanism. This is a construct-validity gap at the center of the paper and needs to be addressed either by re-analyzing the data with a proper diversity measure or by reframing the conclusion to 'number of roles' rather than 'diverse conceptualizations.'","section":"Section 4.2, Table 1"},{"comment":"The statements 'diverse conceptualizations enhance AI adoption' and 'Mental Models of AI directly influence technology adoption decisions' imply a causal direction, but the study is a cross-sectional survey. The correlation between role count and PU/PEU could reflect reverse causality (users who find a tool useful may be motivated to explore and assign more roles to it) or a third variable such as general engagement with AI tools. The paper should soften the causal language and, at minimum, control for the number of AI tools tried and coding experience in a regression or partial correlation analysis, since these variables are already measured and reported in Table 1.","section":"Abstract and Section 5"},{"comment":"The survey sample is a convenience sample of 102 participants recruited from JetBrains' curated list of people who had previously consented to user studies, and the role options were derived from JetBrains' Developer Ecosystem Report. This dual dependence on JetBrains-affiliated channels may limit the representativeness of both the role distribution and the correlations with acceptance. The authors should explicitly discuss this limitation and temper the generalizability claims, or provide evidence that the sample is diverse in terms of tool usage and professional background beyond the reported experience levels.","section":"Section 3"}],"minor_comments":[{"comment":"The table header contains a typo: 'AT tools tried' should be 'AI tools tried'.","section":"Table 1"},{"comment":"The text mentions 'teacher, mentor, senior colleague, or junior colleague' as roles, but the survey options listed in Section 3 include 'teacher' but not 'mentor' (the closest option is 'senior colleague' or 'companion'). Please align the description with the actual survey options.","section":"Section 4.2"},{"comment":"The factor analysis section reports factor loadings but does not report eigenvalues, the proportion of variance explained, or the scree plot itself. Adding these would strengthen the justification for choosing two factors.","section":"Section 4.2"},{"comment":"The sentence '24+ with 16 and more years of experience' is awkwardly phrased; '24 participants with 16 or more years' would be clearer.","section":"Section 4.2"},{"comment":"The reference to the survey DOI is good, but the paper should also state whether the anonymized dataset and analysis scripts are available, since the manuscript says 'available upon request' in Section 4, which is less transparent than the survey instrument.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"This is a borderline case. The empirical work is relevant and the correlations are real, but the central interpretation is not currently supported by the measured construct. The authors could fix this within the manuscript's scope by re-analyzing role endorsement patterns (e.g., number of factors endorsed, entropy, or an interaction term) and by softening the causal language. I would also gently push the authors to be more explicit about the JetBrains recruitment pipeline, as reviewers and readers may otherwise question the independence of the sample. The paper is not fatally flawed, but it needs substantive revision before it can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the two-factor role taxonomy is a useful addition, and the correlations are honest descriptive facts. But the abstract's causal claim is not what the data test. This is a revision, not a rejection.\n\nWhat is actually new: the paper shows that AI role attributions in software engineering group into Support and Expert factors, and that the total number of roles correlates with PU and PEU in a small survey. That specific mapping is not in the papers they cite, and the qualitative 80/20 split between tool and teammate views is a neat, credible summary. Credit where due: they use standard factor analysis, report factor loadings, and give a full correlation table. For a five-page companion paper, that is a legitimate empirical contribution.\n\nNow the soft spots. The strongest concern is construct validity. Table 1 and the abstract equate \"assigning multiple roles\" with \"diverse conceptualizations.\" But the independent variable is just the count of roles selected. A participant can tick four Support roles and zero Expert roles, producing a high count with a narrow, single mental model; another can tick one Support and one Expert role, producing a low count with genuine diversity. The paper never computes a diversity index, cross-category endorsement, or an interaction between the two factors. So the observed correlation with PU and PEU could be driven by general AI enthusiasm, acquiescence, or depth of engagement rather than by mental-model diversity. That is a real gap at the center of the paper's interpretation.\n\nSecond, the causal wording \"enhance AI adoption\" overstates what cross-sectional correlations from 102 participants can support. The authors could soften this without losing the contribution.\n\nThird, the sample is JetBrains volunteers, and the role options come from the JetBrains ecosystem report. That limits generalizability, though the internal correlations remain informative. The data is \"available upon request\" only; for a short empirical paper, depositing anonymized survey responses would strengthen reproducibility. Also minor: no reliability coefficients for the TAM scales, and the negative correlation between the two factors (-0.15) is left unexplained, though it is not significant.\n\nThe citation pattern is fine. Reusing their own prior interview data is legitimate secondary analysis, and the TAM and mental-model references are appropriate.\n\nWho is this for? Researchers working on AI4SE adoption or HCI for coding tools. It is a modest data point, not a field reorgaIt is a modest data point, not a field reorganization. It deserves a serious referee; I would send it to review, with the expectation of revision. A good referee will ask for a diversity measure, softer wording, and either the data or a stronger justification for keeping it closed.","headline":"A useful role taxonomy for AI coding tools, but the paper's central claim that 'diverse conceptualizations enhance adoption' is not supported by the reported measure, which counts roles rather than diversity.","tokens_in":7662,"tokens_out":1872,"would_cite":true,"duration_ms":21254,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Developers who assign multiple roles to an AI coding tool report it as more useful and easier to use, and the paper treats this as a path to wider adoption.","keywords":["AI4SE","mental models","role attribution","technology acceptance model","perceived usefulness","perceived ease of use","developer survey","factor analysis"],"falsifier":"A replication using a representative sample of developers recruited outside an opt-in user-study panel, allowing open-ended role descriptions instead of a fixed list, would falsify the claim if the number of attributed roles showed no positive correlation with perceived usefulness and perceived ease of use.","tokens_in":6653,"feed_emoji":"🤖","tokens_out":5612,"duration_ms":51061,"temperature":0.7,"pith_summary":"This paper asks how developers mentally frame AI-powered development tools and whether that framing predicts adoption. Drawing on 38 interviews and a survey of 102 developers, it identifies two broad mental models—AI as an inanimate tool and AI as a human-like teammate—and shows statistically that the roles developers assign group into Support Roles and Expert Roles. Its central finding is that the more roles a developer assigns to AI, the higher their Perceived Usefulness and Perceived Ease of Use, which the authors read as evidence that diverse conceptualizations support adoption. If true, this gives concrete design and onboarding levers for AI coding tools: help developers see the tool in multiple roles rather than a single narrow one.","feed_headline":"More roles for AI means more adoption, survey finds","feed_subtitle":"Developers who see AI as both tool and teammate rate it higher on usefulness and ease of use.","key_machinery":"The central object is the role attribution itself: the set of roles a developer says AI plays, measured by a thirteen-option survey derived from a yearly developer ecosystem report. The argument is carried by two statistical structures built on those attributions—a two-factor model distinguishing Expert Roles from Support Roles, and Pearson correlations of role counts with Perceived Usefulness and Perceived Ease of Use from a Revised TAM questionnaire. The factor model and the correlation table are what turn qualitative talk about 'assistant' or 'colleague' into a claim about adoption.","core_discovery":"The paper's central claim is that role attribution is not a passive byproduct of using AI tools but a measurable factor in technology acceptance. In the survey data, the total number of roles assigned correlates with both TAM scales ($r = 0.59$ with perceived usefulness, $r = 0.56$ with perceived ease of use, both $p < 0.001$), and both factor-derived role dimensions correlate positively with acceptance. Factor analysis with varimax rotation and a $>0.4$ loading threshold splits the thirteen offered roles into Expert Roles (advisor, reviewer, problem solver) and Support Roles (assistant, reference guide, tool). The authors interpret this as showing that developers who conceptualize AI along multiple dimensions—both helper and expert, both tool and teammate—find it more useful and easier to integrate, and they argue this is consistent with the qualitative divide between tool-minded and teammate-minded developers.","pith_inferences":["The observed correlation may partly reflect a general AI enthusiasm factor: developers who like AI might both assign more roles and rate the tool more favorably, a possibility the correlational design cannot separate.","A testable extension would be to experimentally prime a single role, such as 'junior colleague' versus 'tool', and measure whether actual task performance and sustained usage shift, rather than only self-reported perceptions.","The two mental models may correspond to different trust-calibration strategies: a tool framing may invite verification and reduce over-reliance, while a teammate framing may increase tolerance but risk over-trust—a trade-off the paper leaves implicit.","The interview pattern in which novices described AI as a teacher and experienced developers as a junior engineer suggests a developmental trajectory for role attribution that longitudinal studies could track."],"forward_implications":["If role count predicts perceived usefulness and ease of use, designers can nudge adoption by helping developers see a tool in more than one role, for instance as both assistant and reviewer.","Because Support and Expert role attributions load on separate factors, tools that support both kinds of framing may cover a wider range of developer expectations than tools aimed at a single role.","The qualitative pattern—tool-minded developers enforcing stricter technical standards while teammate-minded developers tolerate imperfections—implies that a single onboarding message will not fit all users.","The positive correlation between role attribution and the number of AI tools tried suggests that exposure and conceptualization may reinforce each other in an adoption cycle."],"supporting_citations":[{"why":"It supplies the 38 interviews re-analyzed here to identify the tool-versus-teammate mental models.","marker":"[15]"},{"why":"It provides the Revised TAM questionnaire whose perceived-usefulness and perceived-ease-of-use scales are the acceptance measures.","marker":"[10]"},{"why":"It supplies the thirteen role options, such as assistant and advisor, that the survey offered to participants.","marker":"[9]"},{"why":"It defines perceived usefulness and perceived ease of use as the determinants of adoption that the study correlates with roles.","marker":"[18]"},{"why":"It establishes the prior link between mental models and human-AI team performance that motivates treating role attribution as consequential.","marker":"[2]"},{"why":"It provides the mental-models concept used to frame the qualitative analysis of developers' descriptions.","marker":"[16]"},{"why":"It sets the factor-loading threshold above which roles were interpreted as belonging to a factor.","marker":"[13]"},{"why":"It supplies the factor-analysis implementation used to extract the two-role factors.","marker":"[3]"}],"fun_headline_variants":["Seeing AI as both tool and teammate boosts adoption","Multiple AI roles linked to higher adoption rates","Diverse AI roles predict better tool adoption","How developers view AI shapes its adoption","Role richness in AI tools drives adoption"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 102 people on the tool maker's opt-in study list, and the thirteen role labels the survey offered them, represent the broader population of developers and the roles developers actually attribute to AI.","fun_headline_variants_meta":{"raw":{"variants":["Seeing AI as both tool and teammate boosts adoption","Multiple AI roles linked to higher adoption rates","Diverse AI roles predict better tool adoption","How developers view AI shapes its adoption","Role richness in AI tools drives adoption"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000493,"raw_usage":{"total_tokens":2368,"prompt_tokens":840,"completion_tokens":1528,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":456,"completion_tokens_details":{"reasoning_tokens":1463}},"tokens_in":456,"tokens_out":1528,"duration_ms":10113,"temperature":1.0,"reasoning_tokens":1463,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:31:45.262564+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication using a representative sample of developers recruited outside an opt-in user-study panel, allowing open-ended role descriptions instead of a fixed list, would falsify the claim if the number of attributed roles showed no positive correlation with perceived usefulness and perceived ease of use.","supporting_citations":[{"cited_title":"(Jim) Lewis","cited_arxiv_id":null,"evidence_quote":"It provides the Revised TAM questionnaire whose perceived-usefulness and perceived-ease-of-use scales are the acceptance measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the thirteen role options, such as assistant and advisor, that the survey offered to participants."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It defines perceived usefulness and perceived ease of use as the determinants of adoption that the study correlates with roles."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It sets the factor-loading threshold above which roles were interpreted as belonging to a factor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the factor-analysis implementation used to extract the two-role factors."}],"review_version":1}