{"id":"3084a7da-9731-4c84-b2b0-9a96f24137f5","arxiv_id":"2607.13798","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In an eight-week state DOT pilot, employees' perceived usefulness of Microsoft 365 Copilot dropped significantly after hands-on use, while most initial champions became less enthusiastic.","lead":"An eight-week pilot of Microsoft 365 Copilot at a state transportation agency found that employees' perceived usefulness of the tool dropped significantly after hands-on use, while ease-of-use, trust, and intention changed little. The study also shows different employee 'personas' moved in opposite directions, with many initial enthusiasts becoming less positive.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pre/post PU comparison rests on unvalidated wording change; no measurement invariance test means the −0.23 decline could be a tense/phrasing artifact.","rationale":"The reader's weakest assumption is precisely the measurement-invariance question. I agree that it is the most load-bearing because the pre/post wording change is explicitly disclosed in §2.4 but never statistically justified. The aggregate PU decline is the anchor of the paper's primary contribution, and all downstream analyses (persona migration, ECM-IT interpretation) are built on the same pre/post scale. A tense shift from 'will improve' to 'improved' is not minimal in psychometric terms: it moves from an evaluative expectation to a factual claim, and respondents may systematically rate actual improvements lower than anticipated ones. Without an invariance test (or at least an item-anchored check), the −0.23 effect could be a pure method effect.\n\nI considered the alternative concern that the number of clusters was selected post hoc (§2.7.2): silhouette and CH indices favored two clusters, but the authors chose three for interpretability and balance. This is a legitimate limitation for the persona structure, but it is not as central: the paper's strongest claim includes the aggregate PU decline, which is independent of clustering, and even a different cluster solution would not change the conclusion that average construct scores changed. Similarly, the lack of public data/code is an accessibility concern, not a correctness threat. The measurement-invariance concern, by contrast, directly threatens the validity of the key quantitative result. The qualitative results and task-use shifts provide triangulation, but they are exploratory keyword counts and do not establish the magnitude of the PU decline. Therefore the concrete test should be a formal invariance analysis or a controlled wording experiment. If scalar invariance fails, the central claim should be downgraded or the data re-analyzed with a correction. Since the paper is currently CONDITIONAL with that caveat, my review supports keeping the verdict conditional rather than rejecting outright, because the concern is addressable with existing data.","tokens_in":21122,"tokens_out":7072,"duration_ms":73706,"concrete_test":"Run multi-group confirmatory factor analysis (MG-CFA) on the item-level data for all four constructs (at minimum PU) across the two waves, testing configural, metric, and scalar invariance (e.g., using lavaan on the raw items, with pre/post as groups). Report fit indices and ΔCFI/ΔRMSEA for nested models. If scalar invariance is not supported (e.g., ΔCFI > 0.01), the composite pre/post scores are not directly comparable and the PU decline may be a wording artifact. As an additional check, re-estimate the Wilcoxon signed-rank test on PU after aligning items with a response-shift model or a within-person anchor; if the effect remains only after invariance is established, the central claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical result—the significant decline in perceived usefulness (PU) after hands-on use (Table 3: −0.23, adjusted p<0.001, r=−0.40)—depends entirely on the comparability of the pre- and post-pilot survey items. Section 2.4 states that the post-pilot instrument 'retained the same construct structure and item content but used minimally revised wording to reflect experienced use rather than anticipated use,' and the items shown in Table 4 are future-tense ('Using Microsoft 365 Copilot will improve my job performance'). No measurement-invariance, differential-item-functioning, or item-level equivalence test is reported. If the past-tense wording is systematically more difficult or more conservative than the future-tense wording, the observed −0.23 decline could be an artifact of the linguistic frame rather than a true recalibration of usefulness perceptions. The same confound is inherited by the persona-migration analysis (§2.7, Table 7), because post-pilot scores are assigned to fixed pre-pilot centroids on the same 1–5 scale; any wording-induced scale shift would alter cluster assignments and could produce the reported 'convergence' and migration percentages even if underlying attitudes were unchanged. The limitations section (§4.5) lists threats (self-report, attrition, context) but does not mention this wording confound, which is arguably the most direct threat to the headline construct change. The reader's conditional verdict flags this as the weakest assumption; it is indeed load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a two-wave matched survey of 124 employees at a state DOT who participated in an eight-week Microsoft 365 Copilot pilot. Perceived usefulness, perceived ease of use, behavioral intention, and trust were measured before training/access and again after eight weeks. Nonparametric tests show a significant aggregate decline in perceived usefulness (−0.23, adjusted p < 0.001, r = −0.40) and small non-significant changes for PEOU, BI, and TR. K-means clustering on the four constructs identifies three baseline personas (Skeptics, Cautiously Positive, Champions), and fixed-centroid assignment is used to track migration, with 40% of Skeptics moving up and 68% of Champions moving down. Secondary analyses examine task-use and concern shifts, and keyword-based content analysis is used to contextualize open-ended responses. The findings are interpreted through TAM and Bhattacherjee's ECM-IT as evidence of expectation recalibration.","tokens_in":21497,"tokens_out":6949,"duration_ms":73714,"significance":"The study addresses a genuine gap: longitudinal evidence on generative AI acceptance in public-sector workforces is scarce. The matched-panel design, attrition check (Table 1), use of nonparametric tests with Benjamini–Hochberg correction, and fixed-centroid migration tracking are methodologically transparent. If the pre/post measures were comparable, the aggregate PU decline and the large individual-level migration would be a useful corrective to one-time acceptance surveys. The paper also reports reliability, response-quality screening, and a specified analysis stack. However, the central inference depends on an untested item-wording change, and the persona results depend on a cluster solution that conflicts with its own selection diagnostics. These are not minor issues: they directly affect the headline and secondary claims.","major_comments":[{"comment":"The post-pilot instrument changed wording from future-oriented expectations to past-oriented experiences, but no measurement-invariance, differential-item-functioning, or item-level equivalence test is reported. The post-pilot items are not shown in the paper, so the magnitude of the wording shift is unverifiable. The headline result in Table 3 (PU −0.23, adjusted p < 0.001) depends entirely on pre/post comparability, making this confound load-bearing. The non-significant changes in PEOU/BI/TR do not rule out construct-specific tense effects. Section 4.5 lists threats but omits this one. The authors need either to provide equivalence evidence or to state plainly that the PU decline cannot be interpreted as uncontaminated evidence of expectation recalibration.","section":"§2.4/Table 4 and §3.2/Table 3"},{"comment":"The Silhouette Score and Calinski–Harabasz Index favored a two-cluster solution, yet the three-cluster solution was retained because the two-cluster solution was imbalanced and the three-cluster solution was more interpretable. No silhouette/CH values, no AIC/BIC comparisons from the GMM robustness check mentioned in §2.7.4, and no sensitivity analysis using k=2 or k=4 are reported. All subsequent migration percentages (Table 7) and path-level interpretations (Table 8) depend on the chosen k. This selection needs quantitative justification and robustness reporting before the persona-migration results can be evaluated.","section":"§2.7.2 and §3.4"},{"comment":"The claim that 'upward movement was associated with gains in PU, BI, and TR' is partly definitional, because post-pilot persona assignment is based on fixed centroids computed from exactly these four constructs. A participant moves from Skeptics to Cautiously Positive precisely when their standardized PU/PEOU/BI/TR vector is closer to the C1 centroid, so increases in these constructs are built into the migration rule. The path-specific deltas should be framed as descriptive consequences of the assignment rule, not as evidence of a separate psychological process. Validation with external variables (task-use, concern items, open-ended content) is needed, or the causal wording should be removed. In addition, paths with n=2 (C2→C0) are overinterpreted despite the note.","section":"§3.5/Table 8 and §4.1"}],"minor_comments":[{"comment":"The post-pilot wording of the items is not included; Table 4 only shows pre-pilot phrasing. An appendix with both forms would help readers assess the comparability concern raised in the major comments.","section":"§2.4"},{"comment":"The Gaussian mixture model robustness check is mentioned but no results (AIC/BIC values or profile comparisons) are reported anywhere in the paper. Either report the results or remove the claim.","section":"§2.7.4"},{"comment":"Several percentage columns sum to 99% or 101% due to rounding (e.g., Table 5 overall row, Table 7 pre-pilot percentages). Please correct or add a rounding note.","section":"Tables 5 and 7"},{"comment":"Figure 2 is referenced, but the baseline LLM-use and Microsoft 365 usage patterns are only described verbally. A brief numerical summary in the text would strengthen reproducibility.","section":"§3.2/Figure 2"},{"comment":"Reference formatting is inconsistent in places (e.g., [7], [14]) with irregular capitalization and access-date formatting. Please harmonize with the journal style.","section":"References"}],"recommendation":"reject","confidential_remarks":"The main barrier is the untested item-wording change. This is not a presentation issue; it undermines the central empirical claim that perceived usefulness declined after hands-on use. Because the post-pilot wording is not provided and no equivalence analysis is reported, the authors cannot currently support the headline result with the existing data. The cluster-selection issue is also substantial. I would be open to a resubmission if the authors can supply measurement-equivalence evidence or substantially reframe the contribution to acknowledge that the PU decline is not identifiable from this design."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a genuinely useful longitudinal dataset and a competently executed applied study, but the headline claim needs a caveat. A matched 124-person two-wave survey of a state DOT's M365 Copilot pilot is rare, and the finding—PU dropped after use while PEOU, BI, and trust didn't—is the kind of evidence public agencies need. The transition matrix is the standout: 68% of Champions moved to less enthusiastic personas while 40% of Skeptics became cautiously positive. That's a real contribution, not a restatement.\n\nWhat the paper does well: attrition check, nonparametric tests with BH correction, reliability, fixed-centroid assignment for tracking, and sensible secondary analyses. The task-use and concern shifts (data/chart and presentation tasks declined; job/skills concerns rose) corroborate the recalibration story independently of the PU items. Citation pattern is fine.\n\nSoft spots, in proportion: the measurement invariance worry is legitimate. The post-pilot items were reworded from future to past tense, and the paper doesn't show them or test equivalence. That doesn't make the PU decline fake—the qualitative and task-use results point the same direction—but it means the exact effect size and p-value shouldn't be treated as clean. This should be fixed or explicitly bounded. The migration-path analysis is partly definitional: clusters are built from PU/PEOU/BI/TR, so saying upward movers gained on those constructs is close to a tautology. The transition rates themselves are not tautological, so the main story survives. Choosing k=3 when silhouette/CH favored 2 is a judgment call; they justify it with interpretability and PCA, which is acceptable but post hoc. No data/code public; scripts “upon request” is a minor reproducibility ding.\n\nWho it's for: researchers and practitioners in public-sector AI adoption. It deserves a serious referee; I'd send it out. With revisions—show the post items, test or bound the wording effect, and reframe migration deltas as descriptive—it would be a solid applied contribution.","headline":"A solid, clearly reported longitudinal adoption study whose headline PU decline is credible but rests on an untested wording change; persona migration is partly definitional but the transition rates are a real contribution.","tokens_in":21949,"tokens_out":3741,"would_cite":true,"duration_ms":42753,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"After eight weeks of hands-on use of Microsoft 365 Copilot at a state transportation department, employees' perceived usefulness fell significantly while ease of use, intention, and trust held steady — yet beneath that stable surface, 40% o","keywords":["generative AI adoption","Technology Acceptance Model","expectation-confirmation theory","perceived usefulness","adoption personas","cluster analysis","public sector","longitudinal survey"],"falsifier":"Re-administer the post-pilot survey to a matched group using the original future-tense pre-pilot wording, or run a measurement-invariance analysis (e.g., multi-group confirmatory factor analysis or item-level differential functioning) on the two waves; if the perceived-usefulness decline shrinks to non-significance when wording is held constant, the paper's central recalibration claim is an artifact of item phrasing. Complementary evidence: objective usage logs (prompt counts, task types, correction rates) would show whether the self-reported decline tracks actual engagement.","tokens_in":21039,"feed_emoji":"📉","tokens_out":6559,"duration_ms":63295,"temperature":0.7,"pith_summary":"This longitudinal study follows 124 state transportation employees through an eight-week Microsoft 365 Copilot pilot, measuring perceived usefulness, perceived ease of use, behavioral intention, and trust before and after hands-on use. The central claim is that actual use recalibrates expectations: perceived usefulness declined significantly (mean change −0.23, adjusted p < 0.001), a pattern the authors read as negative disconfirmation of pre-pilot optimism, while the other three constructs showed only small, non-significant changes. That seemingly calm average conceals large individual movement — k-means clustering identifies three baseline acceptance personas (Skeptics, Cautiously Positive, Champions), and the transition matrix shows 40% of Skeptics moving up to Cautiously Positive while 68% of Champions slid to less enthusiastic personas. The upshot for public agencies is that adoption should be monitored dynamically and supported through persona-specific training, workflow examples, verification routines, and trust-calibration safeguards. A fair reader would care because this is one of the first longitudinal accounts of generative AI acceptance inside a public agency, and it reframes a post-pilot drop in usefulness as recalibration rather than failure.","feed_headline":"AI usefulness scores fell after hands-on use in DOT pilot","feed_subtitle":"Forty percent of skeptics warmed up while 68% of champions cooled; averages hide the churn.","key_machinery":"The central mechanism is a matched two-wave survey design combined with k-means clustering on standardized Technology Acceptance Model-plus-trust composites (perceived usefulness, perceived ease of use, behavioral intention, and trust). Baseline clusters are computed from pre-pilot responses, and post-pilot responses are forced onto the same fixed centroids, producing a transition matrix that tracks individual persona migration rather than just aggregate means. The interpretive lens is expectation-confirmation theory: directional post-use change in usefulness, intention, and trust is read as positive or negative disconfirmation of pre-use expectations.","core_discovery":"The paper claims that in an eight-week enterprise pilot of Microsoft 365 Copilot at a state department of transportation, employees' anticipated usefulness of the tool fell significantly after real use (Wilcoxon signed-rank test, adjusted p < 0.001, r = −0.40), while perceived ease of use, behavioral intention, and trust did not change significantly. Interpreting the shift through expectation-confirmation theory, the authors see negative disconfirmation: pre-use expectations, formed without hands-on experience (81% of participants reported minimal or no prior Copilot use), exceeded confirmed experience. Persona analysis shows three baseline groups — Skeptics, Cautiously Positive, and Champio","pith_inferences":["A testable extension the paper leaves open: link survey responses to objective use telemetry (login frequency, prompt volume, correction rates) to see whether behavioral intention actually predicts sustained, verified use.","If expectation recalibration is real, a six- or twelve-month follow-up wave should show the perceived-usefulness decline plateauing rather than continuing; the paper's own expectation-confirmation framing implies this.","The sharp rise in job-and-skills concerns among downward movers suggests skill anxiety may be a causal driver of migration, not just a correlate; a dedicated measure of perceived AI substitution threat would sharpen the mechanism.","Because post-pilot personas are assigned to fixed pre-pilot centroids, the meaning of 'Champion' is anchored to pre-pilot norms; modeling time-varying clusters could reveal whether the persona labels themselves shift after experience."],"forward_implications":["If the perceived-usefulness decline reflects recalibration rather than failure, post-pilot usefulness scores are a better baseline for forecasting long-term continuance than pre-pilot enthusiasm.","Aggregate acceptance statistics alone understate workforce heterogeneity; transition matrices should be part of enterprise AI rollout monitoring.","Trust behaves as the dynamic load-bearing construct: gains track upward migration and losses track downward migration, so interventions should target calibrated trust rather than generic enthusiasm.","Hands-on experience narrows the task-use portfolio toward communication and summarization and away from data/chart and presentation work, implying workflow-specific training and tool refinement.","Rising job-and-skills concerns alongside falling accuracy and privacy concerns mean governance should address role identity and professional judgment, not only data security."],"fun_headline_variants":["AI usefulness drops after hands-on use in DOT pilot","Expectation recalibration: Copilot usefulness falls after trial","Skeptics warm up but champions cool down in AI pilot","Persona churn: usefulness drops, but 40% of skeptics shift","AI pilot: usefulness wanes as expectations recalibrate"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the pre- and post-pilot surveys measure the same four constructs in the same way. In §2.4 the paper states that post-pilot items used 'minimally revised wording to reflect experienced use rather than anticipated use,' and no measurement-invariance test is reported; if the wording change shifted item meaning or difficulty, the significant perceived-usefulness decline could be an artifact of phrasing rather than genuine expectation recalibration","fun_headline_variants_meta":{"raw":{"variants":["AI usefulness drops after hands-on use in DOT pilot","Expectation recalibration: Copilot usefulness falls after trial","Skeptics warm up but champions cool down in AI pilot","Persona churn: usefulness drops, but 40% of skeptics shift","AI pilot: usefulness wanes as expectations recalibrate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1222,"prompt_tokens":817,"completion_tokens":405,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":318}},"tokens_in":561,"tokens_out":405,"duration_ms":4209,"temperature":1.0,"reasoning_tokens":318,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T03:39:14.766640+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-administer the post-pilot survey to a matched group using the original future-tense pre-pilot wording, or run a measurement-invariance analysis (e.g., multi-group confirmatory factor analysis or item-level differential functioning) on the two waves; if the perceived-usefulness decline shrinks to non-significance when wording is held constant, the paper's central recalibration claim is an artifact of item phrasing. Complementary evidence: objective usage logs (prompt counts, task types, correction rates) would show whether the self-reported decline tracks actual engagement.","supporting_citations":[],"review_version":1}