REVIEW 3 major objections 5 minor 28 references
After eight weeks of hands-on use of Microsoft 365 Copilot at a state transportation department, employees' perceived usefulness fell significantly while ease of use, intention, and trust held steady — yet beneath that stable surface, 40% o
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 03:39 UTC pith:D5XLPTIY
load-bearing objection A solid, clearly reported longitudinal adoption study whose headline PU decline is credible but rests on an untested wording change; persona migration is partly definitional but the transition rates are a real contribution. the 3 major comments →
Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that in an eight-week enterprise pilot of Microsoft 365 Copilot at a state department of transportation, employees' anticipated usefulness of the tool fell significantly after real use (Wilcoxon signed-rank test, adjusted p < 0.001, r = −0.40), while perceived ease of use, behavioral intention, and trust did not change significantly. Interpreting the shift through expectation-confirmation theory, the authors see negative disconfirmation: pre-use expectations, formed without hands-on experience (81% of participants reported minimal or no prior Copilot use), exceeded confirmed experience. Persona analysis shows three baseline groups — Skeptics, Cautiously Positive, and Champio
What carries the argument
The central mechanism is a matched two-wave survey design combined with k-means clustering on standardized Technology Acceptance Model-plus-trust composites (perceived usefulness, perceived ease of use, behavioral intention, and trust). Baseline clusters are computed from pre-pilot responses, and post-pilot responses are forced onto the same fixed centroids, producing a transition matrix that tracks individual persona migration rather than just aggregate means. The interpretive lens is expectation-confirmation theory: directional post-use change in usefulness, intention, and trust is read as positive or negative disconfirmation of pre-use expectations.
Load-bearing premise
The load-bearing premise is that the pre- and post-pilot surveys measure the same four constructs in the same way. In §2.4 the paper states that post-pilot items used 'minimally revised wording to reflect experienced use rather than anticipated use,' and no measurement-invariance test is reported; if the wording change shifted item meaning or difficulty, the significant perceived-usefulness decline could be an artifact of phrasing rather than genuine expectation recalibration
What would settle it
Re-administer the post-pilot survey to a matched group using the original future-tense pre-pilot wording, or run a measurement-invariance analysis (e.g., multi-group confirmatory factor analysis or item-level differential functioning) on the two waves; if the perceived-usefulness decline shrinks to non-significance when wording is held constant, the paper's central recalibration claim is an artifact of item phrasing. Complementary evidence: objective usage logs (prompt counts, task types, correction rates) would show whether the self-reported decline tracks actual engagement.
If this is right
- If the perceived-usefulness decline reflects recalibration rather than failure, post-pilot usefulness scores are a better baseline for forecasting long-term continuance than pre-pilot enthusiasm.
- Aggregate acceptance statistics alone understate workforce heterogeneity; transition matrices should be part of enterprise AI rollout monitoring.
- Trust behaves as the dynamic load-bearing construct: gains track upward migration and losses track downward migration, so interventions should target calibrated trust rather than generic enthusiasm.
- Hands-on experience narrows the task-use portfolio toward communication and summarization and away from data/chart and presentation work, implying workflow-specific training and tool refinement.
- Rising job-and-skills concerns alongside falling accuracy and privacy concerns mean governance should address role identity and professional judgment, not only data security.
Where Pith is reading between the lines
- A testable extension the paper leaves open: link survey responses to objective use telemetry (login frequency, prompt volume, correction rates) to see whether behavioral intention actually predicts sustained, verified use.
- If expectation recalibration is real, a six- or twelve-month follow-up wave should show the perceived-usefulness decline plateauing rather than continuing; the paper's own expectation-confirmation framing implies this.
- The sharp rise in job-and-skills concerns among downward movers suggests skill anxiety may be a causal driver of migration, not just a correlate; a dedicated measure of perceived AI substitution threat would sharpen the mechanism.
- Because post-pilot personas are assigned to fixed pre-pilot centroids, the meaning of 'Champion' is anchored to pre-pilot norms; modeling time-varying clusters could reveal whether the persona labels themselves shift after experience.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a two-wave matched survey of 124 employees at a state DOT who participated in an eight-week Microsoft 365 Copilot pilot. Perceived usefulness, perceived ease of use, behavioral intention, and trust were measured before training/access and again after eight weeks. Nonparametric tests show a significant aggregate decline in perceived usefulness (−0.23, adjusted p < 0.001, r = −0.40) and small non-significant changes for PEOU, BI, and TR. K-means clustering on the four constructs identifies three baseline personas (Skeptics, Cautiously Positive, Champions), and fixed-centroid assignment is used to track migration, with 40% of Skeptics moving up and 68% of Champions moving down. Secondary analyses examine task-use and concern shifts, and keyword-based content analysis is used to contextualize open-ended responses. The findings are interpreted through TAM and Bhattacherjee's ECM-IT as evidence of expectation recalibration.
Significance. The study addresses a genuine gap: longitudinal evidence on generative AI acceptance in public-sector workforces is scarce. The matched-panel design, attrition check (Table 1), use of nonparametric tests with Benjamini–Hochberg correction, and fixed-centroid migration tracking are methodologically transparent. If the pre/post measures were comparable, the aggregate PU decline and the large individual-level migration would be a useful corrective to one-time acceptance surveys. The paper also reports reliability, response-quality screening, and a specified analysis stack. However, the central inference depends on an untested item-wording change, and the persona results depend on a cluster solution that conflicts with its own selection diagnostics. These are not minor issues: they directly affect the headline and secondary claims.
major comments (3)
- [§2.4/Table 4 and §3.2/Table 3] The post-pilot instrument changed wording from future-oriented expectations to past-oriented experiences, but no measurement-invariance, differential-item-functioning, or item-level equivalence test is reported. The post-pilot items are not shown in the paper, so the magnitude of the wording shift is unverifiable. The headline result in Table 3 (PU −0.23, adjusted p < 0.001) depends entirely on pre/post comparability, making this confound load-bearing. The non-significant changes in PEOU/BI/TR do not rule out construct-specific tense effects. Section 4.5 lists threats but omits this one. The authors need either to provide equivalence evidence or to state plainly that the PU decline cannot be interpreted as uncontaminated evidence of expectation recalibration.
- [§2.7.2 and §3.4] The Silhouette Score and Calinski–Harabasz Index favored a two-cluster solution, yet the three-cluster solution was retained because the two-cluster solution was imbalanced and the three-cluster solution was more interpretable. No silhouette/CH values, no AIC/BIC comparisons from the GMM robustness check mentioned in §2.7.4, and no sensitivity analysis using k=2 or k=4 are reported. All subsequent migration percentages (Table 7) and path-level interpretations (Table 8) depend on the chosen k. This selection needs quantitative justification and robustness reporting before the persona-migration results can be evaluated.
- [§3.5/Table 8 and §4.1] The claim that 'upward movement was associated with gains in PU, BI, and TR' is partly definitional, because post-pilot persona assignment is based on fixed centroids computed from exactly these four constructs. A participant moves from Skeptics to Cautiously Positive precisely when their standardized PU/PEOU/BI/TR vector is closer to the C1 centroid, so increases in these constructs are built into the migration rule. The path-specific deltas should be framed as descriptive consequences of the assignment rule, not as evidence of a separate psychological process. Validation with external variables (task-use, concern items, open-ended content) is needed, or the causal wording should be removed. In addition, paths with n=2 (C2→C0) are overinterpreted despite the note.
minor comments (5)
- [§2.4] The post-pilot wording of the items is not included; Table 4 only shows pre-pilot phrasing. An appendix with both forms would help readers assess the comparability concern raised in the major comments.
- [§2.7.4] The Gaussian mixture model robustness check is mentioned but no results (AIC/BIC values or profile comparisons) are reported anywhere in the paper. Either report the results or remove the claim.
- [Tables 5 and 7] Several percentage columns sum to 99% or 101% due to rounding (e.g., Table 5 overall row, Table 7 pre-pilot percentages). Please correct or add a rounding note.
- [§3.2/Figure 2] Figure 2 is referenced, but the baseline LLM-use and Microsoft 365 usage patterns are only described verbally. A brief numerical summary in the text would strengthen reproducibility.
- [References] Reference formatting is inconsistent in places (e.g., [7], [14]) with irregular capitalization and access-date formatting. Please harmonize with the journal style.
Circularity Check
Persona-migration associations partly restate the k-means assignment rule; central PU decline is independent.
specific steps
-
self definitional
[§2.7.3–2.7.4, Table 8; interpreted in §4.1]
"These transformed post-survey responses were then fed into the previously trained K-Means model to determine each participant’s post-survey cluster label. ... For each transition path, we also computed mean changes in PU, PEOU, BI, and TR to describe the construct-level shifts associated with upward, downward, or stable persona movement."
Post-pilot cluster labels are assigned by nearest k-means centroid in the same PU/PEOU/BI/TR space used to define the baseline personas. Because the centroids are ordered from Skeptics (low) to Champions (high), an 'upward' move means the post vector is closer to a higher centroid; a 'downward' move means it is closer to a lower centroid. The change vector therefore necessarily has positive projection on the centroid-difference direction for upward moves and negative projection for downward moves. Reporting mean gains in PU/BI/TR for upward movers and declines for downward movers (Table 8) restates the clustering/assignment rule, with only the magnitudes being empirical.
-
self definitional
[§3.4, Table 6]
"Pairwise post-hoc comparisons were conducted following the significant Kruskal–Wallis tests to identify which personas differed from one another. The results indicated that all three persona pairs were significantly different across each construct at the pre-pilot stage. This confirms that the three baseline personas represented distinct levels of technology acceptance prior to the pilot."
The personas are the output of k-means clustering applied to exactly the four constructs (PU, PEOU, BI, TR) tested in Table 6. K-means partitions observations to maximize between-centroid separation on those variables, so statistically significant between-persona differences on the same variables are an artifact of the clustering objective, not independent evidence that the personas represent distinct acceptance groups. The Kruskal–Wallis and post-hoc tests therefore validate the algorithm's own criterion rather than providing external confirmation of the typology.
full rationale
The central empirical result—the significant decline in perceived usefulness after hands-on use (Table 3: −0.23, adjusted p<0.001, r=−0.40)—is not circular: it comes from a paired pre/post comparison of independently collected Likert-scale items, and the wording change (future- vs past-tense items) is a measurement-validity concern, not a definitional reduction. Likewise, the reported migration counts (40% of Skeptics moving to Cautiously Positive, 68% of Champions moving downward) are empirical frequencies. However, two persona-level claims reduce partly by construction. First, the conclusion that upward migration was associated with gains in usefulness, behavioral intention, and trust is nearly tautological under fixed-centroid assignment: cluster labels are assigned by nearest centroid in the same four-construct space, so moving to a higher persona entails moving toward higher centroid coordinates. Only the magnitudes in Table 8 add empirical content. Second, the paper presents between-persona significance tests on the clustering variables as confirmation of distinct personas, but k-means was fit to those very variables, so the significant Kruskal–Wallis and post-hoc results are forced by the partition criterion. No load-bearing self-citation or ansatz-smuggled-in-via-citation pattern was found; the overlapping-author citations ([14], [19]) are background or methodological and do not carry the argument. Overall, the paper's headline PU decline and migration counts are independent, but a substantial interpretive layer of the persona analysis is definitional, warranting a partial-circularity score rather than a clean pass.
Axiom & Free-Parameter Ledger
free parameters (3)
- Number of clusters k =
3
- K-means centroids (standardized space) =
Not reported numerically in standardized space; Table 5 shows raw-scale cluster means
- Keyword dictionary terms for qualitative analysis =
Not quantified
axioms (7)
- domain assumption TAM and trust constructs measured by four 5-point Likert items each are valid indicators of PU, PEOU, BI, and TR
- domain assumption Pre/post item wording changes do not alter construct meaning
- domain assumption Self-reported perceptions approximate actual attitudes and behavior
- domain assumption Attrition is essentially random and the retained sample is representative
- domain assumption K-means Euclidean geometry on standardized constructs produces meaningful adoption personas
- domain assumption ECM-IT disconfirmation is an appropriate interpretive lens for the aggregate PU decline
- standard math Standard statistical test assumptions for Wilcoxon, McNemar, and Kruskal-Wallis tests
invented entities (2)
-
Acceptance personas (Skeptics, Cautiously Positive, Champions)
no independent evidence
-
Stable/Improved vs Declined transition groups
no independent evidence
read the original abstract
Generative AI tools are increasingly being piloted in public agencies, but limited evidence explains how employee acceptance changes after hands-on use. This study examines Microsoft 365 Copilot adoption during an eight-week pilot at a state Department of Transportation. A matched two-wave survey measured perceived usefulness, perceived ease of use, behavioral intention, and trust before and after participation. After matching and response-quality screening, the sample included 124 employees. Nonparametric tests assessed aggregate changes, k-means clustering identified baseline acceptance personas, and fixed-centroid assignment tracked migration. Open-ended responses were examined using keyword-based content mapping. Perceived usefulness declined significantly after use, suggesting recalibration of expectations, while perceived ease of use, behavioral intention, and trust showed only small, nonsignificant changes. Three baseline personas emerged: Skeptics, Cautiously Positive users, and Champions. Although persona counts changed modestly, individual movement was substantial: 40 percent of Skeptics moved to Cautiously Positive, while 68 percent of Champions moved to less enthusiastic personas. Upward movement was associated with gains in usefulness, behavioral intention, and trust; downward movement was associated with declines in usefulness and trust. Communication and summarization remained stable use cases, while data, chart, and presentation tasks declined. Accuracy and privacy concerns decreased, but job and skills concerns increased. Public-sector AI adoption should be monitored dynamically and supported through persona-specific training, workflow examples, verification routines, and trust-calibration safeguards. The study offers a framework for tracking workforce heterogeneity during enterprise generative AI implementation.
Figures
Reference graph
Works this paper leans on
-
[1]
L. Bies, S. Schmidt, S. Morana, D. Werth, Future office: A comparative study on the acceptance and utilization of generative ai technologies, in: 2024 International Conference on Artificial Intelligence, Computer, Data Sciences and Applications (ACDSA), 2024, pp. 1–6. doi:10.1109/ACDSA59508.2024.10467651
arXiv 2024
-
[2]
URL: https://www.morganstanley.com/ideas/morgan-stanley-tmt-conference-2 023-barcelona, morgan Stanley Ideas blog, accessed 29 July 025
Morgan Stanley, Generative AI’s Impact and Opportunity in Tech and Beyond, 2023. URL: https://www.morganstanley.com/ideas/morgan-stanley-tmt-conference-2 023-barcelona, morgan Stanley Ideas blog, accessed 29 July 025
2023
-
[3]
T. Eloundou, S. Manning, P. Mishkin, D. Rock, Gpts are gpts: An early look at the labor–market impact potential of large language models, arXiv preprint arXiv:2303.10130 (2023). URL:https://arxiv.org/abs/2303.10130
Pith/arXiv arXiv 2023
-
[4]
T. Orchard, L. Tasiemski, The rise of generative ai and possible effects on the economy, Economics and Business Review 9 (2023). URL: http://dx.doi.org/10.18559/ebr. 2023.2.732. doi:10.18559/ebr.2023.2.732
-
[5]
Carrasco, C
M. Carrasco, C. Habib, F. Felden, R. Sargeant, S. Mills, S. Shenton, J. Ingram, G. Dando, Generative AI for the Public Sector: From Opportunities to Value, Technical Report, Boston Consulting Group, Boston, MA, 2024. URL: https://www.bcg.com/publicatio ns/2024/generative-ai-for-the-public-sector-from-opportunities-to-value , white paper
2024
-
[6]
URL: https://www.mckinsey.com/capabilities/quantumblack/our-insights/th e-state-of-ai-in-2023-generative-ais-breakout-year, accessed: 2025-07-29
McKinsey & Company, The state of AI in 2023: Generative AI’s breakout year, 2023. URL: https://www.mckinsey.com/capabilities/quantumblack/our-insights/th e-state-of-ai-in-2023-generative-ais-breakout-year, accessed: 2025-07-29. 29
2023
-
[7]
Mayer, L
H. Mayer, L. Yee, M. Chui, R. Roberts, Superagency in the workplace: Empowering people to unlock AI’s full potential at work, 2025. URL: https://www.mckinsey.com /capabilities/mckinsey-digital/our-insights/superagency-in-the-workplace -empowering-people-to-unlock-ais-full-potential-at-work , mcKinsey Digital report, accessed 29 Jul 2025
2025
-
[8]
now comes the hard part, https://www.microsoft.com/en-us/worklab/work-trend-index/a i-at-work-is-here-now-comes-the-hard-part, 2024
Microsoft, LinkedIn, 2024 work trend index annual report: AI at work is here. now comes the hard part, https://www.microsoft.com/en-us/worklab/work-trend-index/a i-at-work-is-here-now-comes-the-hard-part, 2024. Accessed 31 May 2025
2024
-
[9]
URL: https: //blogs.microsoft.com/blog/2023/03/16/introducing-microsoft-365-copilot -your-copilot-for-work/, accessed: 2025-07-29
Microsoft, Introducing Microsoft 365 Copilot—your copilot for work, 2023. URL: https: //blogs.microsoft.com/blog/2023/03/16/introducing-microsoft-365-copilot -your-copilot-for-work/, accessed: 2025-07-29
2023
-
[10]
N. A. of State Chief Information Officers (NASCIO), M. . Company, Generative artificial intelligence and its impact on state government it workforces, 2024. URL: https: //www.nascio.org/resource-center/resources/generative-artificial-int elligence-and-its-impact-on-state-government-it-workforces/ , accessed 2025-07-29
2024
-
[11]
S. Heo, S. Na, Ready for departure: Factors to adopt large language model (LLM)-based artificial intelligence technology in the AEC industry, Results in Engineering 25 (2025) 104325. doi:10.1016/j.rineng.2025.104325
arXiv 2025
-
[12]
M. A. Beltran, M. I. Ruiz Mondragon, S. H. Han, Comparative analysis of generative ai risks in the public sector, in: Proceedings of the 25th Annual International Conference on Digital Government Research (DGO 2024), 2024, pp. 613–622. doi: 10.1145/3657054. 3657125
doi:10.1145/3657054 2024
- [13]
-
[14]
A. Kumar, A. Hosseini, A. Azarbayjani, A. Heydarian, O. Shoghli, Adoption of AI-assisted e–scooters: The role of perceived trust, safety, and demographic drivers, 2025. URL: https://arxiv.org/abs/2502.05117. arXiv:2502.05117, version 1, submitted 7 Feb 2025
Pith/arXiv arXiv 2025
-
[15]
A. Choudhury, H. Shamszare, Investigating the impact of user trust on the adoption and use of chatgpt: Survey analysis, Journal of Medical Internet Research 25 (2023) e47184. URL:https://www.jmir.org/2023/1/e47184. doi:10.2196/47184, pMID: 37314848
doi:10.2196/47184 2023
-
[16]
F. D. Davis, Perceived usefulness, perceived ease of use, and user acceptance of information technology, MIS Quarterly 13 (1989) 319–340. doi:10.2307/249008
doi:10.2307/249008 1989
-
[17]
P. Shrivastava, Understanding acceptance and resistance toward generative ai technologies: a multi-theoretical framework integrating functional, risk, and sociolegal factors, Frontiers in Artificial Intelligence Volume 8 - 2025 (2025). URL: https://www.frontiersin. 30 org/journals/artificial-intelligence/articles/10.3389/frai.2025.1565927 . doi:10.3389/fr...
arXiv 2025
-
[18]
A. Bhattacherjee, Understanding information systems continuance: An expectation- confirmation model, MIS Quarterly 25 (2001) 351–380. doi:10.2307/3250921
doi:10.2307/3250921 2001
-
[19]
A. Karimzadeh, S. Sabeti, O. Shoghli, Optimal clustering of pavement segments using K-prototype algorithm in a high-dimensional mixed feature space, Journal of Management in Engineering 37 (2021) 04021041. doi:10.1061/(ASCE)ME.1943-5479.0000910
arXiv 2021
-
[20]
A. Bansal, M. Sharma, S. Goel, Improved k-mean clustering algorithm for prediction analysis using classification technique in data mining, International Journal of Computer Applications 157 (2017) 35–40. URL: http://dx.doi.org/10.5120/ijca2017912719 . doi:10.5120/ijca2017912719
-
[21]
T. Kanungo, D. Mount, N. Netanyahu, C. Piatko, R. Silverman, A. Wu, An efficient k-means clustering algorithm: analysis and implementation, IEEE Transactions on Pattern Analysis and Machine Intelligence 24 (2002) 881–892. URL: http://dx.doi.o rg/10.1109/tpami.2002.1017616. doi:10.1109/tpami.2002.1017616
Pith/arXiv arXiv 2002
-
[22]
Venkatesh, F
V. Venkatesh, F. D. Davis, A theoretical extension of the technology acceptance model: Four longitudinal field studies, Management science 46 (2000) 186–204
2000
-
[23]
Venkatesh, M
V. Venkatesh, M. G. Morris, G. B. Davis, F. D. Davis, User acceptance of information technology: Toward a unified view1, MIS quarterly 27 (2003) 425–478
2003
-
[24]
Gefen, E
D. Gefen, E. Karahanna, D. W. Straub, Trust and tam in online shopping: An integrated model1, MIS quarterly 27 (2003) 51–90
2003
-
[25]
D. H. Mcknight, M. Carter, J. B. Thatcher, P. F. Clay, Trust in a specific technology: An investigation of its components and measures, ACM Transactions on management information systems (TMIS) 2 (2011) 1–25
2011
- [26]
-
[27]
F. Dell’Acqua, E. McFowland III, E. R. Mollick, H. Lifshitz-Assaf, K. Kellogg, S. Ra- jendran, L. Krayer, F. Candelon, K. R. Lakhani, Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality, Working Paper 24-013, Harvard Business School, 2023. doi:10.2139/ssrn.4573321
-
[28]
OMB Memorandum M–24–10
Office of Management and Budget, Advancing governance, innovation, and risk manage- ment for agency use of artificial intelligence, 2024. OMB Memorandum M–24–10. 31
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.