REVIEW 5 major objections 5 minor 18 references
The World of AI: A Novel Approach to AI Literacy for First-year Engineering Students
T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read An interdisciplinary, math-free AI course for first-year engineering students significantly improves AI literacy on all five measured rubric domains.
desk verdict A useful AI-literacy curriculum write-up whose headline causal claim is undercut by a pre/post design that pairs individual survey scores with group presentation scores, plus a required/elective contradiction. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the five-category presentation rubric—Understanding of Future Problem, Future Relevance, AI's Contribution to the Problem, Corrective AI Solution, and Responsibility & Safeguards—used as the instrument for measuring learning outcomes. The mechanism is the mapping of pre-course survey questions onto those same five rubric categories, which makes it possible to compare baseline written answers with final group presentations and to run paired Wilcoxon signed-rank tests with Bonferroni correction. The survey items themselves were adapted from established instruments, including the Meta AI Literacy Scale and UNESCO's AI Ethics Readiness Assessment, which gives the measurement its claimed construct validity.
What would settle it
Re-score the anonymized pre-course open-ended responses using the same five-category rubric that the external evaluator used for final presentations, with raters blind to which answers came from before the course; if the pre-course scores are not lower than the presentation scores, the reported gains would be an artifact of measurement. A complementary test would give the identical instrument (same questions, same scoring) to a control group that does not take the course and check whether the gains replicate.
Extended reading notes
Core claim
The discovery is that a 14-week, three-module course—AI and the Planet, AI and the Society, AI and the Workplace—co-delivered by engineering and humanities faculty moves first-year engineering students' AI literacy from baseline to advanced on a five-category rubric. Scores rose from 5.0–7.0 pre-course to 8.2–9.0 post-course on a 10-point scale, with Bonferroni-adjusted Wilcoxon $p$-values from $6.03\times10^{-13}$ to $1.16\times10^{-22}$ and rank-biserial effect sizes $r=.47$ to $.77$. The paper claims all five improvements are both statistically robust and educationally meaningful, with four rubrics showing large practical gains and one showing a moderate–large gain.
Load-bearing premise
The pre-course survey questions and the final presentation rubric are treated as measuring the same five constructs, so that a rise in scores means learning rather than a difference in task difficulty, individual versus group work, or scoring method.
Editorial extensions
If this is right
- Other engineering programs could adopt the three-module structure and see similar literacy gains if the course design is the cause.
- AI ethics and societal impact can be taught to first-year students without prerequisite mathematics or computer science coursework.
- The co-teaching model pairing engineering and humanities faculty provides a template for other technical-ethics courses.
- The largest effect sizes on Corrective AI Solution and Responsibility & Safeguards suggest the course is especially effective at building ethical reasoning and solution design.
- The result supports the case for making such a course a required first-semester gateway rather than an upper-level elective.
Reading between the lines
- Because pre-course scores came from individual written answers and post-course scores from group presentations rated by an external evaluator, part of the gap could reflect task format rather than learning; a direct comparison would use identical instruments at both time points.
- The LLM-based scoring of qualitative responses introduces a potential confound, since an automated model may reward length or style rather than content mastery.
- Without a control group, the observed gains could partly reflect maturation or exposure to other coursework; a randomized or waiting-list design would separate these factors.
- The claim would be strengthened by a follow-up semester re-administering the same survey to test whether the gains persist.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes the design, delivery, and evaluation of 'The World of AI,' a one-semester interdisciplinary AI literacy course for first-year engineering students at Plaksha University. The course combines three modules (AI and the Planet, AI and the Society, AI and the Workplace) with co-taught engineering and humanities sessions, guest lectures, and a final group presentation project. Evaluation is based on a single-group pre-post design with 168 students: a pre-course survey (multiple-choice and open-ended items) and final group presentation scores rated on a five-criterion rubric, plus a post-course survey of self-reported perception shifts. The authors report statistically significant pre-to-post gains on all five rubric domains using paired Wilcoxon signed-rank tests with Bonferroni-adjusted p-values and rank-biserial effect sizes, and interpret the results as affirmatively answering the research question about the course's impact on students' understanding of AI's role in the planet, society, and workplace.
Significance. If the reported effects were credible, the paper would be a useful contribution to the growing literature on AI literacy education in undergraduate engineering, particularly for its interdisciplinary design, its focus on societal and ethical dimensions, and its attempt to measure learning across multiple constructs. The course design is described in sufficient detail to be replicable, the survey instruments are adapted from established scales (MAILS and UNESCO's AI Ethics Readiness Assessment), ethical approval is reported, and the use of Bonferroni corrections and effect sizes is a welcome step beyond simple p-value reporting. The acknowledged limitations (absence of a control group, reliance on self-report) are appropriate. However, the central quantitative claim rests on a statistical comparison whose validity is undermined by the unit-of-analysis mismatch between individual pre-survey scores and shared group presentation scores, and by an unvalidated mapping between qualitatively different measurement instruments.
major comments (5)
- [Section 4 (Results) and Section 3 (Data Collection)] The paired Wilcoxon signed-rank tests are reported with n=168, but the endpoint measure is the rubric-based score of final group presentations, in which 'each group envisioned a future problem' (Section 3). If a single presentation score was assigned to all group members, the 168 pairs are not independent and the test statistics, Bonferroni-adjusted p-values, and rank-biserial effect sizes are invalid, because the effective sample size is the number of groups, not the number of students. This problem is logically prior to the construct-comparability issue: even if the pre-survey and the rubric measured identical constructs, the paired test would still be invalid under clustering. The manuscript must report the number of groups and group sizes, and either analyze at the group level (e.g., pre-survey scores aggregated within groups compared with group presentation scores) or provide evidence that final presentations were scored individually for each student.
- [Section 4 (Results)] The claim that pre-survey questions were 'organized into five categories corresponding to the final project rubric... allowing for direct comparisons' is not supported by any evidence of measurement equivalence. The baseline consists of individual written answers to multiple-choice and open-ended questions, scored partly by an automated LLM method, while the endpoint is a group oral presentation scored by an external human evaluator on a 0-10 rubric. Differences in task modality, individual versus group performance, scoring method, and social facilitation during group work could each produce gains that have nothing to do with learning. The manuscript should either use the same instrument before and after the course, or provide explicit evidence (e.g., pilot data, inter-rater agreement on the mapping, or a construct-validation argument) that the two measures are directly comparable.
- [Section 3 (Preliminary Data Analysis)] The LLM-based scoring of qualitative responses is described only as 'an automated LLM-based method that compared each answer to an ideal response, normalizing similarity scores to a 10-point scale, with a sample of responses spot-checked by evaluators.' This is not a validated measurement protocol. The manuscript does not specify which LLM was used, what prompt or similarity metric was employed, how the normalization was computed, how the 'ideal response' was constructed, or what the spot-check agreement was. Because the open-ended qualitative responses feed directly into several rubric categories that are later used in the headline pre-post comparisons, the reliability and validity of this scoring method is load-bearing. The authors should report inter-rater reliability (e.g., Cohen's kappa or ICC between LLM and human evaluators) and provide the full scoring protocol or a supplemental example.
- [Section 4 (Results), Rubrics 3 and 5] For two of the five rubric domains, the pre-post gain is measured with multiple-choice questions that closely mirror content directly taught in the course: Rubric 3 asks about data-center energy consumption and environmental impact, and Rubric 5 asks about workplace bias domains and equitable data collection for underrepresented communities. Such items assess recall of course-specific facts rather than generalizable AI literacy, so the reported gains may partly reflect teaching to the assessment. The authors should re-analyze the data after excluding these content-recall items, or supplement the analysis with transfer tasks that require applying the concepts to novel scenarios, and should temper the claim that all five learning outcomes improved in a construct-valid manner.
- [Section 5 (Discussion)] The Discussion states that the research question is 'affirmatively answered' and that the course caused the observed improvements, yet it simultaneously acknowledges the absence of a control group and reliance on self-reported data. A single-group pre-post design cannot rule out maturation, history, testing effects, or other confounding influences, and the self-report survey asks students to rate how much their understanding changed, which is circular as evidence of learning. The conclusion should be reframed as evidence of pre-post gains within one cohort, with all causal language removed, unless a comparison group or another identification strategy is added.
minor comments (5)
- [Section 3 (Data Collection)] The rubric description is internally inconsistent: it first says each parameter is rated on a '0-10 scale' and later says 'each criterion was rated on a scale from 1 (poor) to 10 (excellent).' Please clarify whether the scale is 0-10 or 1-10.
- [Section 4 (Results), Figure 2] Figure 2's caption states the scale is '1 = Not at all, 10 = Completely,' but the Preliminary Data Analysis section describes Likert responses scored 1-5 and then scaled to 10. Please explain the exact scaling used for the post-course survey items and whether the figure's values are raw or scaled.
- [Section 3 (Data Collection)] The manuscript does not report the number of groups, the average group size, or whether final presentations were scored individually or per group; this information is needed to assess the statistical analyses and also for replication.
- [Section 3 (Course Design)] There is a typo in 'This structure facilitated active engagement, consistent feedbacks' — it should be 'consistent feedback.' Other small language issues (e.g., 'The course comprises of') should be corrected during copyediting.
- [Section 3 (Data Collection)] The manuscript states that instruments were adapted from MAILS and UNESCO's AI Ethics Readiness Assessment but does not describe which items were adapted, how they were modified, or whether the adapted versions were piloted or validated for this population. A brief description or appendix would strengthen the construct-validity argument.
Circularity Check
No significant circularity: the course evaluation compares pre-survey responses to externally scored final presentations; the alignment of survey items to rubric categories is a design choice, not an equation-level reduction.
full rationale
The paper is an empirical course-evaluation study rather than a derivation or prediction from first principles. No fitted parameter is later called a prediction, no uniqueness theorem is imported from the authors' prior work, and there is no self-citation chain: the reference list contains no self-citations. The central pre/post comparison uses a baseline pre-course survey and an endpoint consisting of final group presentations scored by an external evaluator on a five-category rubric. The paper's statement that pre-survey questions were 'organized into five categories corresponding to the final project rubric... allowing for direct comparisons' is an alignment of measurement constructs, not a definitional identity that forces the reported gains; the pre-survey answers and the presentation scores are distinct measurements obtained through different tasks and scored by different means. The post-course survey items ask students to rate the extent of change in their understanding, which is self-report and therefore weak evidence, but the paper explicitly acknowledges reliance on self-reported data and the absence of a control group as limitations; this is a validity caveat, not a circular derivation. The reviewer's concerns about clustered group scores (shared final-presentation scores assigned to all group members) and about teaching-to-the-assessment are statistical and construct-validity threats, not instances where a claimed result is equivalent to its inputs by construction. Accordingly, no circular step meeting the specified evidentiary bar can be quoted, and the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Equal rubric weights across five criteria =
0.2 each
- Likert scale scoring values =
5 to 1 for Strongly Agree to Strongly Disagree
- LLM similarity normalization =
Similarity scores normalized to 0-10
assumptions (4)
- domain assumption Single-group pre-post scores can estimate the course's effect on understanding
- domain assumption Pre-survey open-ended responses and final presentation rubric scores measure the same constructs
- domain assumption LLM-based qualitative scoring approximates human expert judgment
- domain assumption Self-reported perception change reflects actual understanding
Cite this review
Pith. "Pith review of The World of AI: A Novel Approach to AI Literacy for First-year Engineering Students." pith.science (2026). https://pith.science/paper/RGK5Z3ZT
@misc{pith2026250608041,
author = {Pith},
title = {Pith review of: The World of AI: A Novel Approach to AI Literacy for First-year Engineering Students},
year = {2026},
howpublished = {\url{https://pith.science/paper/RGK5Z3ZT}},
note = {Machine review of arXiv:2506.08041}
}
read the original abstract
This work presents a novel course titled The World of AI designed for first-year undergraduate engineering students with little to no prior exposure to AI. The central problem addressed by this course is that engineering students often lack foundational knowledge of AI and its broader societal implications at the outset of their academic journeys. We believe the way to address this gap is to design and deliver an interdisciplinary course that can a) be accessed by first-year undergraduate engineering students across any domain, b) enable them to understand the basic workings of AI systems sans mathematics, and c) make them appreciate AI's far-reaching implications on our lives. The course was divided into three modules co-delivered by faculty from both engineering and humanities. The planetary module explored AI's dual role as both a catalyst for sustainability and a contributor to environmental challenges. The societal impact module focused on AI biases and concerns around privacy and fairness. Lastly, the workplace module highlighted AI-driven job displacement, emphasizing the importance of adaptation. The novelty of this course lies in its interdisciplinary curriculum design and pedagogical approach, which combines technical instruction with societal discourse. Results revealed that students' comprehension of AI challenges improved across diverse metrics like (a) increased awareness of AI's environmental impact, and (b) efficient corrective solutions for AI fairness. Furthermore, it also indicated the evolution in students' perception of AI's transformative impact on our lives.
Figures
Reference graph
Works this paper leans on
-
[1]
Journal of Economic Perspectives29(3), 3–30 (2015)
Autor, D.H.: Why are there still so many jobs? the history and future of workplace automation. Journal of Economic Perspectives29(3), 3–30 (2015)
work page 2015
-
[2]
Baker, T., Smith, L.: Educ-ai-tion rebooted? exploring the future of artificial in- telligence in schools and colleges. Tech. rep., NESTA (2019)
work page 2019
-
[3]
arXiv preprint arXiv:2310.06269 (2023), https://arxiv.org/abs/2310.06269 8 S
Feffer, M., Martelaro, N., Heidari, H.: The ai incident database as an edu- cational tool to raise awareness of ai harms: A classroom exploration of effi- cacy, limitations, & future improvements. arXiv preprint arXiv:2310.06269 (2023), https://arxiv.org/abs/2310.06269 8 S. Siddharth et al
arXiv 2023
-
[4]
Frey, C.B., Osborne, M.A.: The future of employment: How susceptible are jobs to computerisation? Technological Forecasting and Social Change114, 254–280 (2017)
work page 2017
-
[5]
In: Proceedings of the 2024 ASEE Annual Conference & Exposition
Goldenkoff, E., Cech, E.A.: Left on their own: Confronting absences of ai ethics training among engineering master’s students. In: Proceedings of the 2024 ASEE Annual Conference & Exposition. American Society for Engineering Education (2024). https://doi.org/10.18260/1-2–47724
doi:10.18260/1-2 2024
-
[6]
ACM Computing Surveys54(6), 1–35 (2021)
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., Galstyan, A.: A survey on bias and fairness in machine learning. ACM Computing Surveys54(6), 1–35 (2021)
2021
-
[7]
Science366(6464), 447–453 (October 2019)
Obermeyer, Z., Powers, B., Vogeli, C., Mullainathan, S.: Dissecting racial bias in an algorithm used to manage the health of populations. Science366(6464), 447–453 (October 2019). https://doi.org/10.1126/science.aax2342
-
[8]
Proceedings of the AAAI Conference on Artificial Intel- ligence37(13), 15834–15842 (2024)
Orchard, A., Radke, D.: An analysis of engineering students’ responses to an ai ethics scenario. Proceedings of the AAAI Conference on Artificial Intel- ligence37(13), 15834–15842 (2024). https://doi.org/10.1609/aaai.v37i13.26880, https://ojs.aaai.org/index.php/AAAI/article/view/26880
Show all 18 references
-
[9]
arXiv preprint arXiv:2104.10350 (2021), https://arxiv.org/abs/2104.10350
Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.M., Rothchild, D., So, D., Texier, M., Dean, J.: Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350 (2021), https://arxiv.org/abs/2104.10350
2021 arXiv
-
[10]
Pillsbury Law (2022), https://www.pillsburylaw.com/en/news- and-insights/ai-impact-vulnerable-workforce.html
Pillsbury, W.S.P.L.: The impact of artificial intelligence on vulnerable populations in the workforce. Pillsbury Law (2022), https://www.pillsburylaw.com/en/news- and-insights/ai-impact-vulnerable-workforce.html
2022
-
[11]
https://www.coursera.org/learn/ethics-of-artificial-intelligence (2024), massive Open Online Course on Coursera, accessed 5 May 2025
Politecnico di Milano: Ethics of artificial intelligence. https://www.coursera.org/learn/ethics-of-artificial-intelligence (2024), massive Open Online Course on Coursera, accessed 5 May 2025
2024
-
[12]
Education Sciences11(7) (2021)
Stadelmann, T., Keuzenkamp, J., Grabner, H., Würsch, C.: The ai-atlas: Didactics for teaching ai and machine learning on-site, online, and hy- brid. Education Sciences11(7) (2021). https://doi.org/10.3390/educsci11070318, https://www.mdpi.com/2227-7102/11/7/318
2021 doi
-
[13]
Computers in Human Behavior: Artificial Humans1(2), 100014 (2023)
Straka, S., Latoschik, M.E., Wienrich, C.: MAILS — meta AI literacy scale: Development and testing of an AI literacy questionnaire based on well- founded competency models and psychological change- and meta-competencies. Computers in Human Behavior: Artificial Humans1(2), 1000...
2023
-
[14]
In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
Strubell, E., Ganesh, A., McCallum, A.: Energy and policy considerations for deep learning in nlp. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. pp. 3645–3650 (2019)
2019
-
[15]
Tzu Chi Medical Journal32(4), 339–343 (2020)
Tai, M.C.T.: The impact of artificial intelligence on human soci- ety and bioethics. Tzu Chi Medical Journal32(4), 339–343 (2020). https://doi.org/10.4103/tcmj.tcmj_71_20
2020 doi
-
[16]
UNESCO: Readiness assessment methodology: A tool of the recommen- dation on the ethics of artificial intelligence. Tech. rep., United Na- tions Educational, Scientific and Cultural Organization, Paris (2023), https://www.unesco.org/en/articles/readiness-assessment-methodology-...
2023
-
[17]
https://web.stanford.edu/class/cs182/ (2025), course syllabus, accessed 5 May 2025
University, S.: Cs/ethicsoc 182: Ethics, public policy, and technological change. https://web.stanford.edu/class/cs182/ (2025), course syllabus, accessed 5 May 2025
2025
-
[18]
International Journal of STEM Education 11(3), 45–60 (2024)
Usher, M., Barak, M.: Integrating ai ethics into science and engineering curricula: A case for explicit-reflective learning. International Journal of STEM Education 11(3), 45–60 (2024). https://doi.org/10.1186/s40594-024-00493-4
2024 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.