REVIEW 3 major objections 5 minor 7 references
Understanding Student Acceptance, Trust, and Attitudes Toward AI-Generated Images for Educational Purposes
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that undergraduate computer science and software engineering students readily accept and trust AI-generated images for educational tasks, with failures in realism and text rendering as the main limit on detail-oriented…
desk verdict A modest, methodically described exploratory study whose abstract overstates trust: the reported means are moderate, not high. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a task-based usability protocol combined with three standardized instruments. Fifteen undergraduate participants used ChatGPT 4.0 through a desktop browser to generate images from ten pre-defined prompts spanning themes such as technology, culture, nature, and health; they then completed the Technology Acceptance Model questionnaire (adapted to measure self-efficacy, social norms, perceived usefulness, and behavioral intention), the Trust Scale, and the Technology Attitude Scale, followed by open-ended interviews. The Likert-scale scores provide the quantitative picture of acceptance, trust, and attitudes, while thematic analysis of the interviews supplies the realism and accuracy concerns that explain why trust stays only moderate.
What would settle it
A replication with a larger, randomly selected sample of undergraduate computer science and software engineering students at several universities would settle the claim: if the mean acceptance, trust, and attitude scores fall at or below the neutral midpoints of their scales, the reported high acceptance, trust, and positive attitudes do not generalize.
Extended reading notes
Core claim
The central claim, stated the way the author would state it, is that AI-generated images are already an accepted and trusted aid for many undergraduate educational tasks, but they are not yet a precision tool. Students rated perceived usefulness and ease of use well above the scale midpoint (for example, TA11 M=4.1, TA13 M=4.5), reported strong confidence and enjoyment (TAS2 M=4.6, TAS3 M=4.5), and expressed clear intention to keep using the images in future coursework. Trust was more guarded, with competence, dependability, and reliability ratings in the 3.3-3.7 range. Interview responses explain the gap: images are praised as beautiful and immediately usable for creative and presentation work, but criticized for unrealistic human depictions, incorrect text, and failure to follow detailed prompts, making them unsuitable for technical diagrams and other accuracy-critical assignments. The paper takes this combination as evidence that students will adopt AI images now for creative and illustrative purposes, while more demanding educational uses await improvements in realism, prompt fidelity, and access.
Load-bearing premise
The load-bearing assumption is that fifteen self-selected volunteers recruited through university social media groups represent undergraduate computer science and software engineering students broadly enough for the study's general statements about student acceptance, trust, and attitudes.
Editorial extensions
If this is right
- AI-generated images can serve as ready-to-use visual content for presentations, reports, and creative design tasks, where students judge them as immediately usable.
- Realism and text-rendering failures will keep AI-generated images out of technical diagrams and other accuracy-critical assignments until generators can follow prompts with higher fidelity.
- Cost and access are practical adoption barriers, so affordable or institutionally provided access to AI image tools should increase their educational use.
- The coexistence of high acceptance with moderate trust suggests that educational guidelines addressing accuracy, intellectual property, and quality standards are needed before AI images become routine in coursework.
Reading between the lines
- A testable extension beyond the paper is to compare disciplines: the precision complaint would predict lower usefulness ratings in engineering, medicine, and science than in design and humanities courses.
- Because trust scores lag behind attitude scores, one visible high-stakes failure—a technical diagram with wrong labels in a graded report—could suppress trust more sharply than the overall positive attitudes imply.
- The realism and text-rendering limitations point to a curricular use the paper leaves implicit: treating AI image generators as draft tools whose outputs students must verify and edit rather than adopt as finished.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports an exploratory mixed-methods study of 15 undergraduate computer science and software engineering students at one Thai university. Participants used ChatGPT 4.0 to generate images from ten pre-defined prompts, then completed adapted Technology Acceptance Model (TAM), Trust Scale, and Technology Attitude Scale (TAS) questionnaires and a structured interview. The paper claims that students show high acceptance, trust, and positive attitudes toward AI-generated images for educational purposes, while concerns about realism and technical precision limit use in detail-oriented tasks.
Significance. The study addresses a genuine gap: student perspectives on AI-generated images in higher education are underrepresented relative to technical and instructor-focused work. The mixed-methods design, the use of adapted published instruments, and the thematically organized qualitative findings are strengths. If the central descriptive claims were properly calibrated, the study would provide a useful exploratory baseline for future research. However, the current quantitative evidence does not support the 'high trust' headline, and the sample is too narrow and self-selected to support broad generalizations about 'students'.
major comments (3)
- [Section 4.3 and Abstract] The abstract's claim of 'high trust' is not supported by the data reported in Section 4.3. The overall trust mean is 3.3 (SD=0.7), with dimension means of 3.3 to 3.7 on a 5-point scale, i.e., close to the neutral midpoint. No inferential test against a neutral baseline, confidence interval, or comparison standard is provided, so the data are equally consistent with moderate or neutral trust. The Discussion correctly softens to 'moderate to high trust' (Section 5), but the Abstract and Conclusion (Section 6) revert to unqualified positive claims. This is load-bearing because the paper's stated contribution is the empirical demonstration of high acceptance, trust, and positive attitudes. I recommend either reporting baseline comparisons (e.g., one-sample tests against the midpoint, with appropriate caution for n=15) or systematically replacing 'high' with 'moderate' or 'mixed' throughout the abstract, discussion, and conclusion.
- [Section 3 (Method), participant recruitment] The study's external validity claims are not supported by the sampling procedure. Fifteen volunteers recruited via one university's social media student groups from two programs cannot support statements about 'students' in general, as made in the Abstract and Conclusion. If the sample overrepresents students already interested in AI, all reported means could be inflated. The paper should reframe all general conclusions as specific to this sample and treat the study as hypothesis-generating, or provide a sampling strategy and justification of representativeness if broader claims are intended.
- [Section 3 (Method), instruments and procedure] The measurement of 'positive attitudes' rests on the 'Technology Attitude Scale' attributed to Rosen et al. (2013), but that reference describes the Media and Technology Usage and Attitudes Scale, which is a broader instrument; the paper does not specify which items or subscales were adapted, nor does it include the actual items. Additionally, the text states that the entire study required about 45 minutes per participant, while Table 2 sums to 70 minutes (10 pre-study + 30 in-study + 30 post-study). Both issues affect reproducibility: readers cannot tell what was measured or how long the procedure actually took.
minor comments (5)
- [Section 4.5.6] The phrase 'The Participants emphasized' contains an unnecessary capital 'P' in 'Participants' and should be corrected.
- [References] Creswell (2014) is listed in the references but never cited in the text; Zhu et al. (2021) is cited in Section 1 for ethical challenges, but the listed reference is the CycleGAN paper, which does not address ethics; Epstein et al. (2023) is missing its title in the reference list.
- [Table 1] The theme names in Table 1 use inconsistent punctuation, with some themes ending in a colon; please format the table consistently.
- [Section 4.2] The combined reporting 'TA7, TA8: M = 3.9' obscures which items are being combined; report each item separately or explain the aggregation.
- [Section 4.1 and Figure 1] Figure 1 is not described in the text beyond its caption; add a sentence in Section 4.1 referring to it so readers understand what is shown.
Circularity Check
No significant circularity: the study reports empirical survey and interview data without deriving predictions from fitted inputs or self-cited constraints.
full rationale
The paper is an exploratory usability study that collects Likert-scale questionnaire responses and interview feedback from 15 undergraduate students. Its claims about acceptance, trust, and attitudes are direct summaries of the reported means (e.g., TA11: M=4.1, SD=0.8; Trust: M=3.3 to 3.7; TAS: M=4.1 to 4.7) and thematic interview excerpts. No model is fitted, no parameter is calibrated to a target result, and no outcome variable is defined in terms of the explanatory variables. The only cited sources for instruments are prior published scales (Davis, 1989; Merritt, 2011; Rosen et al., 2013; Aburbeian et al., 2022), and the paper does not invoke any self-citation or uniqueness theorem to force its conclusions. The gap between the abstract's 'high trust' wording and the trust-scale means near the scale midpoint is an internal interpretation or accuracy issue, not a circularity. The derivation chain, such as it is, runs from raw responses to descriptive statistics and qualitative themes, with no step that reduces to its own inputs.
Assumptions & free parameters
assumptions (4)
- domain assumption The adapted TAM, Trust Scale, and Technology Attitude Scale measure the intended constructs (acceptance, trust, attitude) in this population.
- domain assumption The fifteen self-selected participants from software engineering and IT programs are treated as an adequate basis for exploratory claims about student attitudes.
- domain assumption Participants' self-reports in questionnaires and interviews reflect their true perceptions and behavior.
- domain assumption The thematic analysis of interview transcripts is reliable and unbiased.
Cite this review
Pith. "Pith review of Understanding Student Acceptance, Trust, and Attitudes Toward AI-Generated Images for Educational Purposes." pith.science (2026). https://pith.science/paper/AQNAYPTP
@misc{pith2026241115710,
author = {Pith},
title = {Pith review of: Understanding Student Acceptance, Trust, and Attitudes Toward AI-Generated Images for Educational Purposes},
year = {2026},
howpublished = {\url{https://pith.science/paper/AQNAYPTP}},
note = {Machine review of arXiv:2411.15710}
}
read the original abstract
Recent advancements in artificial intelligence (AI) have broadened the applicability of AI-generated images across various sectors, including the creative industry and design. However, their utilization in educational contexts, particularly among undergraduate students in computer science and software engineering, remains underexplored. This study adopts an exploratory approach, employing questionnaires and interviews, to assess students' acceptance, trust, and positive attitudes towards AI-generated images for educational tasks such as presentations, reports, and web design. The results reveal high acceptance, trust, and positive attitudes among students who value the ease of use and potential academic benefits. However, concerns regarding the lack of technical precision, where the AI fails to accurately produce images as specified by prompts, moderately impact their practical application in detail-oriented educational tasks. These findings suggest a need for developing comprehensive guidelines that address ethical considerations and intellectual property issues, while also setting quality standards for AI-generated images to enhance their educational use. Enhancing the capabilities of AI tools to meet precise user specifications could foster creativity and improve educational outcomes in technical disciplines.
Figures
Reference graph
Works this paper leans on
-
[1]
Introduction Recent advancements in artificial intelligence (AI), such as deep learning and neural networks (Liang et al., 2022), have significantly impacted various sectors, including education. Notably, a research study by Fitria (2021) highlights how these technological improvements have boosted the efficiency of both teaching and learning, benefiting ...
work page 2021
-
[2]
Related Studies The article by Rubman (2024) in the MIT Sloan Review classifies AI -generated images in educational contexts into three distinct categories: instructive, decorative, and distracting. It emphasizes that these images can significantly enhance learning when strategically chosen to align with educational goals an d foster inclusivity. Further,...
work page 2024
-
[3]
Method The usability testing was conducted in a university lab setting with 15 undergraduate students from the Software Engineering and Information Technology programs, all aged between 18 and 25. Recruitment was facilitated through the university's social media student groups, with participants meeting specific inclusion criteria: current enrollment in t...
work page 2022
-
[4]
Perceived Quality of AI -Generated Images
Results and findings 4.1. Pre-study questions The study involved 15 students aged 18 -25, with a gender distribution of 53.33% female (8) and 46.67% male (7). The use of AI -generated images for academic purposes varied among participants: 33.3% used them moderately, 26.7% rarely, 20% infrequently, and 20% frequently. No participants reported very frequen...
-
[5]
Discussion In this study, the findings reveal complex yet interesting responses from the participants, characterized by robust engagement and notable recognition, alongside identified challenges and opportunities for future enhancements. The findings show that s tudents demonstrate d high self -efficacy, using AI tools independently, which is crucial for ...
work page 2021
-
[6]
Conclusion In conclusion, this study highlights the strong engagement and self -efficacy of students with AI - generated images in educational contexts, showing high acceptance of the technology. While positive attitudes and trust are prevalent, concerns over realism, accuracy, and technical limitations suggest areas for enhancement. The potential of AI -...
-
[7]
References Aburbeian, A. M., Owda, A. Y., & Owda, M. (2022). A technology acceptance model survey of the metaverse prospects. AI, 3(2), 285 -302. https://doi.org/10.3390/ai3020018 Adeshola, I., & Adepoju, A. P. (2023). The opportunities and challenges of ChatGPT in education. Interactive Learning Environments. https://doi.org/10.1080/10494820.2023.2253858...
arXiv 2022
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.