{"id":"84afc33c-adf1-460c-83fc-5b706bfbfa07","arxiv_id":"2507.05046","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Frequent use, perceived expertise, and ethical risk perceptions predict university students' trust in ChatGPT, while trust varies strongly by task type.","lead":"A survey of 115 university students and four follow-up interviews found that regular use, perceived competence, and ethical worries predict trust in ChatGPT more than demographics. The findings give universities and product teams concrete areas to target for safer, better-calibrated use of generative AI.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Risk-dimension sign ambiguity, hidden in a missing Appendix A, makes the 'ethical risk is a strongest predictor' headline uninterpretable; the abstract also overstates non-significant ease/transparency effects.","rationale":"The reader's verdict identified the missing measurement appendix as the core weakness, and my reading agrees. I would sharpen it: the 'risk' dimension is not just unreported, it is signed ambiguously in the text itself (Figure 2 vs. Discussion), and the headline that expertise and ethical risk are the strongest predictors depends entirely on how the Risk composite is reverse-coded. That makes it the single most load-bearing point, since a one-line scoring change could flip the sign of the second strongest predictor. The paper's other patterns (usage frequency, task contingency, societal attitudes) are less affected because they rely on single items or straightforward scales. I also note that the abstract's 'secondary effects' for ease of use and transparency is not supported by Table 5's p-values, but this is a framing fix and does not change the conditional verdict. The appropriate outcome remains CONDITIONAL: the claims can be checked and should be once Appendix A and item-level data are supplied, and the authors should relabel or clarify the Risk dimension direction.","tokens_in":16776,"tokens_out":7168,"duration_ms":86019,"concrete_test":"Retrieve Appendix A and, using the raw item scores and reverse-coding rule, reconstruct the seven composites; then recompute the Spearman correlations in Figure 2 and the multiple regression in Table 5 under both possible codings of the Risk items (risk-as-perceived and risk-as-compliance). If the sign or significance of Risk changes between codings, or if neither coding reproduces both the reported Figure 2 correlations and the positive beta in Table 5, then the abstract's claim that expertise and ethical risk are the strongest predictors is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing issue is the scoring and direction of the 'risk' dimension, which underpins the claim that ethical risk is one of the two strongest predictors (Table 5). Section 3.2 calls this dimension 'risk (framed as ethical compliance)' and says reverse coding was 'applied where necessary', but the items and coding are deferred to Appendix A, which is absent. The internal evidence is contradictory: Figure 2 shows Risk positively correlated with Ease of Use (r = 0.43, p = .047) and with the other positive dimensions, which is coherent only if a high score means low perceived risk / high ethical compliance; yet the Discussion states 'Ease of use was inversely related to perceived risk', treating high Risk as high perceived risk. Table 5 also shows a positive Risk coefficient. Without knowing the item wording and reverse-coding direction, the reader cannot tell whether the reported association means 'more concern about risk predicts more trust' or 'better ethical compliance predicts more trust'. If the Risk composite was scored in the opposite direction, the headline ranking of expertise and ethical risk inverts or vanishes. This concern is not merely a missing appendix: it is an internal inconsistency in the sign of the central predictor. A secondary reporting issue is that Table 5 shows ease of use p=.311 and transparency p=.123, so the abstract's 'secondary effects' wording is not supported by the paper's own regression.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a mixed-methods study of trust in ChatGPT among 115 UK university students, combining a survey with four semi-structured interviews. It addresses four research questions: user attributes, seven trust dimensions (expertise, predictability, transparency, human-likeness, ease of use, risk, reputation), task-specific trust, and perceived societal impact. The main quantitative findings are that usage frequency is positively associated with trust while self-reported understanding of LLM mechanics is negatively associated; that perceived expertise and risk are the strongest regression predictors of overall trust; that trust is highest for summarising and coding and lowest for entertainment and citation generation; and that positive societal-impact perceptions are associated with higher trust. The qualitative interviews are used to illustrate and contextualise these patterns. The central claim, as stated in the abstract, is that trust in ChatGPT among university students is primarily shaped by hands-on experience, perceived competence, ethical-risk judgement, and task verififiability.","tokens_in":17051,"tokens_out":4386,"duration_ms":51969,"significance":"If the results hold, the paper makes a useful empirical contribution to the HCI and AI-trust literature by providing a task-level, mixed-methods account of trust in a widely used LLM, and by connecting the seven-dimension framework of Choudhury and Shamszare to a student population. The explicit RQ structure, the use of non-parametric tests appropriate to Likert data, and the inclusion of detailed correlation and regression tables are strengths, as is the Discussion's acknowledgement that the trust-use relationship may be bidirectional. However, the manuscript is not currently verifiable in its headline claims because the questionnaire items, reverse-coding rules, and reliability statistics for the trust-dimension composites are deferred to an appendix that is not included in the arXiv text, and because the direction of the Risk dimension is internally inconsistent. The abstract also overstates the evidence for 'secondary effects' of ease of use and transparency. These issues are fixable, but they are load-bearing for the paper's central ranking of trust dimensions.","major_comments":[{"comment":"The composite trust-dimension scales, their item wording, reverse-coding rules, and reliability statistics are described only as 'provided in Appendix A', which is not included in the arXiv manuscript. All RQ2 results, including the headline ranking of expertise and risk as the strongest predictors in Table 5, rest on these composites actually measuring the intended constructs. The authors must supply Appendix A (or report the items and Cronbach's alpha in the main text) so that readers can verify the scaling, the reverse-coding, and the reliability of each dimension.","section":"Section 2.2.1 / Section 3.2 / Appendix A"},{"comment":"The direction of the Risk dimension is internally contradictory. In Figure 2, Risk correlates positively with Ease of Use (r = 0.43, p = .047) and with the other positive trust dimensions, which is coherent only if a high Risk score means low perceived risk / high ethical compliance. Yet Section 4 states 'Ease of use was inversely related to perceived risk', which treats a high Risk score as high perceived risk, and Table 5 reports a positive Risk coefficient in the regression predicting overall trust. Without the item wording and the reverse-coding direction, the reader cannot determine whether the reported association means 'more concern about risk predicts more trust' or 'better ethical compliance predicts more trust'. If the Risk composite were scored in the opposite direction, the headline ranking of expertise and ethical risk would invert or disappear. This needs to be resolved explicitly.","section":"Figure 2 / Table 5 / Section 4 (Trust Dimensions)"},{"comment":"The abstract states that ease of use and transparency had 'secondary effects' on overall trust, but Table 5 reports p = .311 for ease of use and p = .123 for transparency, so neither is statistically significant in the multiple regression. The text should either describe these dimensions as non-significant predictors in the regression while noting their significant bivariate correlations in Table 4, or support the 'secondary effects' claim with an appropriate analysis such as relative-importance or dominance analysis.","section":"Abstract / Table 5"},{"comment":"The causal wording 'frequent use increased trust' is not supported by the cross-sectional survey design. The Discussion itself acknowledges a possible reciprocal relationship ('It remains conceivable... that trust itself motivates continued use'), which is inconsistent with the causal language used in the abstract and in Section 3.1. Replace causal formulations with associational wording throughout, or explicitly frame the causal interpretation as a hypothesis requiring longitudinal or experimental data.","section":"Section 3.1 / Section 4 (Factors Influencing User Trust)"}],"minor_comments":[{"comment":"The questionnaire is said to be provided in Appendix B, but Appendix B is also not included in the arXiv text. Please include both appendices or state where they can be obtained.","section":"Section 2.2.1 / Appendix B"},{"comment":"The caption says 'Spearman correlation matrix' but the diagonal entries are labelled 'Pearson r' and p-values are shown as p=0.000; use p < .001 and reconcile the correlation-type label.","section":"Figure 2 caption"},{"comment":"The sentence beginning 'Pairwise Dunn tests' is incomplete and is immediately repeated in the following paragraph; the duplicate sentence should be removed and the first completed.","section":"Section 3.4"},{"comment":"The task label 'Editing' in Table 8 should be consistent with 'Proofreading or editing' used in Figure 4 and elsewhere.","section":"Section 3.3 / Table 8"},{"comment":"The variable 'LLM understanding' is self-reported rather than objectively measured; the text should consistently describe it as perceived understanding to avoid overclaiming.","section":"Section 3.1 / Table 2"},{"comment":"Because several predictors are strongly correlated (e.g., Expertise and Predictability, r = 0.74), the regression would benefit from reporting variance inflation factors or standardised coefficients to show the stability of the coefficient ranking.","section":"Table 5 / Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of the journal and the research design is appropriate for a cs.HC audience. I am recommending major revision rather than rejection because the central claims are plausible and the identified problems are fixable: the authors need to provide the missing appendices, clarify the coding and direction of the Risk dimension, and align the abstract with the reported significance levels. I would ask the editor to require the supplementary materials as part of the revision, since the current version cannot be fully evaluated without them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — this one is worth a look, but only with a red pen. The core finding — students' confidence in ChatGPT's citation ability, despite its known inaccuracies, is the single strongest correlate of global trust — is a nice piece of evidence for automation bias. The study also does a decent job applying an existing seven-dimension trust framework to a new population, and the task-level results (high trust for coding/summarising, low for references/entertainment) fit the verification-cost story. The interviews are a plus; they give the numbers some texture.\n\nThe soft spots are real. First, Appendix A, which contains the actual Likert items and reliability stats for the seven dimensions, is missing. That alone would be a minor fix, but the stress-test found a sign ambiguity in the 'risk' dimension: the correlation matrix shows risk positively correlated with ease of use and other positive dimensions, which makes sense only if a high score means 'low risk / high ethical compliance'. The discussion then says 'ease of use was inversely related to perceived risk', which implies the opposite scoring. Without seeing the items and reverse-coding, the headline ranking of 'expertise and ethical risk' as the strongest predictors is uninterpretable. That's a load-bearing problem, not a cosmetic one.\n\nSecond, the abstract says ease of use and transparency had 'secondary effects', but Table 5 shows p=.311 and p=.123. Those are not effects; they're null results. That's an overstatement that should be corrected.\n\nThird, there are some statistical red flags. The correlation matrix in Figure 2 lists p=.032 for an r of .74 with n=115; that p-value is off by many orders of magnitude. Either the reported p-values are from something else or there's a typo. It shakes confidence in the surrounding numbers. Also, the sample is 77 CS students out of 115, so the 'university students' framing is really 'CS-heavy Edinburgh students'.\n\nThe causal wording ('frequent use increased trust') bothers me less because the authors do acknowledge the possible reverse direction in one sentence, but the abstract presents it as causal. A cross-sectional survey can't support that.\n\nBottom line: the descriptive findings are plausible and the citation-confidence result is worth publishing. But the paper needs major revisions before it's ready for a serious journal: include the appendices, resolve the risk-scoring direction, correct the abstract, and clean up the correlation p-values. I'd send it to peer review — the core material is there — but I'd expect a substantial revision. I wouldn't cite it in its current form.","headline":"A useful exploratory study with a striking automation-bias finding, but the missing measurement appendix and a sign ambiguity in the 'risk' dimension make the headline predictor ranking uninterpretable as written.","tokens_in":17555,"tokens_out":3201,"would_cite":false,"duration_ms":36466,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"University students trust ChatGPT when it seems expert and ethically safe, not when it seems human-like.","keywords":["ChatGPT","user trust","trust dimensions","automation bias","task-specific trust","university students","mixed-methods","AI ethics"],"falsifier":"Give readers the full item texts and reliability statistics from the missing appendix and run a confirmatory factor analysis on the seven composite scales; if the items do not separate into seven internally consistent dimensions, the regression ranking of expertise and ethical risk as strongest predictors collapses.","tokens_in":16585,"feed_emoji":"🤖","tokens_out":7382,"duration_ms":74536,"temperature":0.7,"pith_summary":"This mixed-methods study tries to establish what actually drives university students' trust in ChatGPT, using a survey of 115 UK students and four interviews. Its central claim is that trust is a calibrated, task-specific judgement rather than a general attitude: students trust ChatGPT for coding and summarising, whose outputs are easy to check, and distrust it for citation generation and entertainment, where verification is hard or stakes are unclear. Across seven trust dimensions, perceived expertise and ethical risk are the strongest predictors of overall trust; ease of use and transparency matter secondarily, while human-likeness and reputation do not predict trust. The paper also argues that direct experience outweighs demographics: frequent use raises trust, while self-reported understanding of how LLMs work lowers it. The significance, if true, is that trust in generative AI can be shaped by task design, transparency features, and user education rather than by making systems more human-like.","feed_headline":"ChatGPT trust hinges on task and expertise, not human-likeness","feed_subtitle":"Survey of 115 university students: perceived expertise and task verifiability drive trust; anthropomorphism does not.","key_machinery":"The central object is the seven-dimension trust framework adopted from prior work, which decomposes trust into expertise, predictability, transparency, human-likeness, ease of use, ethical risk, and reputation; the study measures each with composite Likert scales and enters all seven in a regression predicting overall trust. The second load-bearing mechanism is task verifiability: users trust outputs they can check, such as code and summaries, and withhold trust where output is hard to verify or high-stakes, such as references and entertainment. The argument runs through these two devices plus an automation-bias lens, in which fluent, confident output is mistaken for factual reliability, shown by citation confidence being the strongest correlate of global trust despite documented inaccuracy.","core_discovery":"The discovery the paper argues for is that student trust in ChatGPT is primarily a function of task verifiability, perceived competence, ethical risk judgement, and hands-on experience. In the regression on all seven trust dimensions, perceived expertise and ethical risk carry the strongest weight, with ease of use and transparency as secondary predictors; human-likeness and reputation are non-significant. Trust is highest for summarising and coding and lowest for entertainment and sourcing references, yet confidence in ChatGPT's referencing ability is the single strongest correlate of overall trust even though the paper notes those citations are often invented, a pattern the authors read as automation bias. Behaviourally, frequent use predicts higher trust while self-reported technical understanding predicts lower trust, and computer-science students only exceed other students in trusting the system for proofreading and writing. The paper takes these results to show that trust is learned through interaction and calibrated by task demands, not conferred by anthropomorphism or reputation.","pith_inferences":["If task verifiability is the underlying mechanism, then interface changes such as showing confidence scores or adding one-click source verification should shift trust in predictable ways; this is a testable design extension the paper does not itself propose.","The negative link between technical understanding and trust may be partly a selection effect rather than a causal effect of knowledge; a longitudinal AI-literacy course with a control group would separate education from pre-existing disposition.","The automation-bias reading implies that students who express high confidence in ChatGPT's referencing may check citations least; logging real citation-checking behaviour would test whether stated trust tracks actual verification.","Because the sample is a single UK university with a computer-science-heavy skew, the relative weights of the seven dimensions are likely to shift in other populations, and the framework needs cross-validation before being treated as a general model."],"forward_implications":["Designers who want appropriate trust should invest in competence signals, transparency, and accuracy cues rather than human-like personas.","Task-level trust ratings imply that LLM features should make verifiability visible: code and summary outputs earn trust, while citation generation needs disclaimers or verification tools.","Because self-reported technical understanding lowers trust, AI-literacy education is a plausible lever for calibrating trust and countering automation bias.","Because frequent use raises trust, repeated positive interactions may build trust, though the paper notes the relationship between trust and use could be bidirectional."],"supporting_citations":[{"why":"Supplies the seven-dimension trust framework that structures RQ2 and the regression on overall trust.","marker":"[37]"},{"why":"Provides the dispositional-situational-learned trust model used to frame trust as dynamic and experience-based.","marker":"[3]"},{"why":"Establishes the automation-reliance research that underlies the paper's automation-bias interpretation.","marker":"[12]"},{"why":"Supports the claim that transparency modulates trust and that fluent outputs can be mistaken for reliability.","marker":"[13]"},{"why":"Supplies the trust-triangulation model used to explain why users verify outputs against other sources.","marker":"[36]"},{"why":"Provides the guidance on calibrated engagement and transparency that the discussion applies to generative AI.","marker":"[42]"},{"why":"Supports the claim that reliability and source pedigree shape operator reliance on decision aids.","marker":"[46]"},{"why":"Supplies the systematic-review definition of automation bias used to label the citation-confidence result.","marker":"[56]"}],"fun_headline_variants":["Trust in ChatGPT rests on task and expertise, not human-likeness","Students trust ChatGPT for coding, not citations, yet overrate its references","Perceived expertise and ethics, not anthropomorphism, drive ChatGPT trust","Automation bias marks ChatGPT trust: referencing confidence despite flaws","Task-contingent trust in ChatGPT: highest for coding, lowest for citations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole ranking of trust dimensions relies on the unpublished questionnaire items actually measuring the seven dimensions they claim to measure, with the reverse-coded risk items scored correctly.","fun_headline_variants_meta":{"raw":{"variants":["Trust in ChatGPT rests on task and expertise, not human-likeness","Students trust ChatGPT for coding, not citations, yet overrate its references","Perceived expertise and ethics, not anthropomorphism, drive ChatGPT trust","Automation bias marks ChatGPT trust: referencing confidence despite flaws","Task-contingent trust in ChatGPT: highest for coding, lowest for citations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000264,"raw_usage":{"total_tokens":1620,"prompt_tokens":981,"completion_tokens":639,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":546}},"tokens_in":597,"tokens_out":639,"duration_ms":7231,"temperature":1.0,"reasoning_tokens":546,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:33:18.956696+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give readers the full item texts and reliability statistics from the missing appendix and run a confirmatory factor analysis on the seven composite scales; if the items do not separate into seven internally consistent dimensions, the regression ranking of expertise and ethical risk as strongest predictors collapses.","supporting_citations":[{"cited_title":"Choudhury and H","cited_arxiv_id":null,"evidence_quote":"Supplies the seven-dimension trust framework that structures RQ2 and the regression on overall trust."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the dispositional-situational-learned trust model used to frame trust as dynamic and experience-based."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the automation-reliance research that underlies the paper's automation-bias interpretation."},{"cited_title":"Zerilli, U","cited_arxiv_id":null,"evidence_quote":"Supports the claim that transparency modulates trust and that fluent outputs can be mistaken for reliability."},{"cited_title":"Rowley and F","cited_arxiv_id":null,"evidence_quote":"Supplies the trust-triangulation model used to explain why users verify outputs against other sources."},{"cited_title":"Artificial intelligence risk management frame- work: Generative artificial intelligence profile (nist ai 600-1)","cited_arxiv_id":null,"evidence_quote":"Provides the guidance on calibrated engagement and transparency that the discussion applies to generative AI."},{"cited_title":"Madhavan and D","cited_arxiv_id":null,"evidence_quote":"Supports the claim that reliability and source pedigree shape operator reliance on decision aids."},{"cited_title":"Goddard, A","cited_arxiv_id":null,"evidence_quote":"Supplies the systematic-review definition of automation bias used to label the citation-confidence result."}],"review_version":1}