{"id":"c125421b-cb9c-449e-b0b2-f782bb6e9e8c","arxiv_id":"2501.07392","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Participants in a 14-week AI literacy course for a broad university audience self-reported gains on all ten AI literacy survey questions, but no objective or independent measure of learning was used.","lead":"The University of Texas at Austin built a one-credit, online AI literacy course for a broad university audience, and the paper reports on its design, survey results, and lessons learned. The report is a useful template for other universities, but its main evidence of learning is self-reported and lacks a control group.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Retrospective pre/post gains may be response-shift artifacts; the paper's unanalyzed weekly 'prior understanding' ratings offer a within-study check that would settle it.","rationale":"The reader's verdict correctly identifies the retrospective pre-post self-report as the weakest link. I agree, and I sharpen it with a specific internal data source: the weekly reflections included 'prior understanding' ratings that could directly test for response-shift bias. The paper cites Howard and Dailey (1979) yet does not use this contemporaneous measure to validate the retrospective baseline. This matters because the ten retrospective items are the only quantitative evidence for the course's primary objective. A secondary but real issue is the sample composition (70% undergraduates, 54% Natural Sciences, 145/788 respondents) and the contradictory enrollment figures (788 sign-ups vs 131+584 in the summary), which further weaken generalizability. The paper's abstract carefully says 'attendees reported gains,' while the Lessons Learned section states participants 'improved their AI literacy,' a causal claim not supported by the design. Conditional acceptance is appropriate if the authors reframe the claim as self-reported gains; the course's qualitative feedback, enrollment breadth, and open materials are genuine contributions. Hence no change to the reader's verdict.","tokens_in":8355,"tokens_out":4255,"duration_ms":40397,"concrete_test":"Compare the aggregate 'prior understanding' ratings from the weekly reflection surveys (collected during the course) with the retrospective 'before the course' ratings on the ten final-survey items, after mapping to comparable scales and subsetting to respondents who completed both. If the contemporaneous prior-understanding mean is significantly higher than the retrospective 'before' mean, response-shift bias is confirmed and the learning-gain claim must be downgraded to 'self-reported retrospective gains'; if they match, the retrospective measure gains credibility.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the course improved attendees' AI literacy rests entirely on the final retrospective pre/post survey (Figure 2). In this design, participants rated their skill 'before the course' after having completed it, so their 'before' ratings are recollections calibrated against their new knowledge—the classic response-shift bias cited by the authors themselves (Howard and Dailey 1979). The claimed average gains of +0.97 to +1.37 and their p-values only show that retrospective 'before' and 'now' ratings differ; they do not establish actual literacy change. The paper does, however, collect a potentially relevant contemporaneous measure: each weekly reflection asked participants to rate their 'prior understanding' of the topic. If these weekly prior-understanding ratings (gathered during the course) were higher than the retrospective 'before course' ratings on comparable scales, that would directly evidence response-shift bias. The authors never report this comparison, so the strongest evidence for the course's primary objective remains vulnerable to a well-documented self-report artifact. Additionally, the survey sample tilts heavily toward undergraduates (70%) and Natural Sciences (54%), and only 145 of 788 enrollees responded, so the positive results may reflect selection. The phrase in Lessons Learned that the audience 'improved their AI literacy' overstates what the design can support; the abstract's 'reported gains' is the defensible phrasing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on the design, delivery, and retrospective evaluation of a one-credit, online AI literacy course offered at the University of Texas at Austin in Fall 2023 to a broad audience of students, faculty, staff, and community members. It describes the course structure, the interdisciplinary lecture schedule, enrollment demographics, weekly reflection data, and a final course survey. The key quantitative evidence consists of ten retrospective pre/post self-assessment items in which participants rated themselves higher 'now' than they recalled 'before' the course, with average gains between +0.97 and +1.37 on a five-point Likert scale, each reported as statistically significant at p < 0.01. The authors conclude that the course improved participants' AI literacy and use the feedback to design a subsequent three-credit course.","tokens_in":8552,"tokens_out":2616,"duration_ms":25996,"significance":"If the effectiveness claims were treated only as participants' self-reported impressions, the paper provides a useful and detailed blueprint for a broadly accessible AI literacy course, including the lecture schedule, institutional support structures, and openly available course materials. The enrollment data across all UT Austin colleges and the candid reporting of challenges with readings and audience heterogeneity are valuable for instructors designing similar courses. The paper is not a rigorous effectiveness study: it uses no control group, no objective literacy measure, and its sole outcome measure is a retrospective self-assessment collected after the intervention. The authors are appropriately cautious in the abstract ('reported gains') but overstate the conclusion in Lessons Learned. The paper's main contribution is as a course design and lessons-learned narrative, and its conclusions should be scaled back accordingly.","major_comments":[{"comment":"The central claim that the course improved participants' AI literacy rests entirely on retrospective pre/post self-ratings, in which the 'before' ratings were collected at the end of the course. The authors cite Howard and Dailey (1979) and Geldhof et al. (2018), both of which are foundational references for response-shift bias, yet the paper never addresses this threat. Response-shift bias predicts exactly the observed pattern: participants recalibrate their recollection of their prior knowledge after being exposed to course content. Without a contemporaneous pre-course questionnaire, a control group, or an objective knowledge test, the gains of +0.97 to +1.37 show only that retrospective recollections differ from current self-assessments, not that literacy changed. The Lessons Learned statement that 'the audience that participated in the final course survey improved their AI literacy' is therefore not supported by the evidence presented; the abstract's 'reported gains' is the defensible phrasing.","section":"Overall Course Survey, Figure 2"},{"comment":"The weekly reflections included a rating of 'prior understanding' of each week's topic, collected contemporaneously before or during the course. This provides a within-study check on response-shift bias that the authors do not report: if the average of the weekly prior-understanding ratings is systematically lower than the retrospective 'before course' ratings on comparable questions, that would directly evidence response-shift. Conversely, if the weekly prior-understanding ratings are similar to or higher than the retrospective 'before' ratings, the retrospective gains would be more credible. The authors should report this comparison, or explain why the two measures are not comparable. As it stands, the only internal validity check available in the data is omitted.","section":"Response to Weekly Surveys"},{"comment":"The survey sample is 145 respondents out of 788 enrollees, and the authors report that 70 percent were undergraduates and 54 percent were affiliated with the College of Natural Sciences. This is a heavily self-selected subsample that overrepresents the most engaged and technically oriented participants. The paper does not report response rates by enrollment category (students versus auditors, or by college), nor does it discuss how selection might bias the reported retrospective gains. Given that auditors were found to be less engaged than students, the absence of this analysis weakens the generalization of the reported improvements to the full enrolled population and should be addressed explicitly.","section":"Overall Course Survey, enrollment demographics"}],"minor_comments":[{"comment":"The word 'assesments' appears twice in the discussion of Williams (2023) and should be corrected to 'assessments'.","section":"Related Work"},{"comment":"The entry 'Elargethical Datasets' appears to be a typographical error for 'larger ethical datasets' and should be corrected.","section":"Table 1"},{"comment":"The word 'asyncrhonous' should be 'asynchronous'.","section":"Future Plans"},{"comment":"The phrase 'aimed at abroad audience' should be 'aimed at a broad audience'.","section":"Relation to Previous Work"},{"comment":"The text states that gains are 'reported' in parentheses following each question, but the term should be 'in parentheses'; additionally, Figure 2 would be more informative with error bars, per-item p-values, and a statement of the statistical test used (e.g., paired t-test or Wilcoxon signed-rank).","section":"Overall Course Survey, Figure 2"},{"comment":"The survey description reports n = 151 for all questions except Q5 with n = 150; the paper should explain the single missing response for Q5.","section":"Overall Course Survey"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is best understood as an experience report on course design rather than a controlled evaluation of learning outcomes. The authors' own citations of the response-shift literature make the methodological gap particularly conspicuous, and the discrepancy between the cautious abstract and the stronger Lessons Learned claim is likely to draw criticism from reviewers and readers. A revision that consistently frames the results as self-reported gains, adds the available within-study check from the weekly prior-understanding ratings, and explicitly discusses response-shift and selection limitations would be publishable as a practice-oriented contribution. I do not see a load-bearing error that would require rejection, but the central claim needs to be reworded and the analysis supplemented."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe paper is a straightforward experience report on a one-credit AI literacy course offered to a broad university audience. Read it for the course design and the institutional context; don't read it for evidence that the course improved AI literacy, because that evidence isn't there.\n\nWhat the paper does well: it gives a clear picture of the course—14 weekly lectures, interdisciplinary speakers, readings, quizzes, weekly reflections, and a final survey. The enrollment data showing all 17 colleges represented and the breakdown by role is useful. The qualitative coding of open-ended responses is reasonable, and the authors are transparent about several limitations, including the lack of baselines for the detailed evaluation. The abstract is carefully worded: \"attendees reported gains.\" That is true.\n\nThe soft spot is the central evaluation claim. The ten AI-literacy questions are retrospective pre/post: participants rated their skill \"before\" after already completing the course. That design invites response-shift bias, and the authors cite Howard and Dailey but don't address it. The stress-test note is right that the weekly reflections asked for a contemporaneous \"prior understanding\" rating each week. If those weekly ratings were higher than the retrospective \"before\" ratings, that would be direct evidence of response-shift. The paper doesn't report that comparison, even though it collected the data. That's a missed opportunity and an actionable fix.\n\nThe sample is also narrow: 145 or 151 respondents out of 788 enrollees, with 70% undergraduate and 54% from Natural Sciences. The positive results could easily be selection. There is a small internal inconsistency—the summary says 131 students and 584 auditors while the enrollment section says 132 undergraduates, 631 auditors, and 25 external participants; the survey text says 145 participants while Figure 2 says n=151. Minor, but it should be cleaned up.\n\nThe bigger issue is language. In \"Lessons Learned\" the authors write that the audience \"improved their AI literacy.\" That overstates what the design can support. The abstract's phrasing, \"reported gains,\" is the right one, and the lessons-learned section should match it.\n\nBottom line: this is a useful case study for instructors building similar courses, with honest discussion of practical challenges. It deserves a serious referee, but the referee should require a revised claim about learning gains and a response to the response-shift check. I'd bring it to a reading group mainly to discuss retrospective pre/post measurement in education.","headline":"Readable case study of a broad-audience AI literacy course; the only evidence for learning gains is retrospective self-report, and a within-study check for response-shift was collected but never reported.","tokens_in":9119,"tokens_out":2846,"would_cite":false,"duration_ms":27172,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A ten-question retrospective survey found that participants rated their AI literacy higher after the course on every item, with gains from 0.97 to 1.37 points on a five-point scale.","keywords":["AI literacy","course design","higher education","retrospective pre-post survey","interdisciplinary teaching","large language models","seminar course","program evaluation"],"falsifier":"Conduct the same ten-item instrument as a true pre-test at the start of the course and again at the end, with a no-course control group; if the course group shows no larger gain than the control, or if the true pre-test ratings already match the retrospective 'before' ratings, the reported improvement is largely a measurement artifact.","tokens_in":8132,"feed_emoji":"🎓","tokens_out":6719,"duration_ms":58696,"temperature":0.7,"pith_summary":"This paper reports on the rapid design and delivery of a one-credit online AI literacy course open to all members of a large public university, including students, faculty, staff, and community members. The authors claim that participants' AI literacy improved: on a ten-question retrospective pre/post survey, every measure rose by between 0.97 and 1.37 points on a five-point Likert scale, with each gain statistically significant at p < 0.01. The course paired lectures on AI fundamentals with interdisciplinary talks on societal impacts, and weekly reflections plus a final survey guided the design of a follow-up three-credit version. The significance is a tested template for bringing AI literacy to non-technical audiences quickly and for gathering actionable feedback that can shape later offerings.","feed_headline":"Ten AI-literacy measures all rose after one course","feed_subtitle":"A one-credit, 14-week seminar for students, staff, and community reports gains up to 1.37 points on a 5-point scale.","key_machinery":"The central mechanism is the retrospective pre/post self-assessment: a ten-item Likert questionnaire administered at the end of the course in which participants rate their AI literacy both before and after the course in a single sitting. This design is intended to control for response-shift bias, in which a participant's internal standard for 'literate' changes as they learn, and the authors explicitly ground the method in the response-shift literature. The survey, together with weekly reflection prompts and thematically coded open-ended responses, carries the argument that the course improved literacy and identifies the design lessons.","core_discovery":"The course's central finding is that a broad-audience AI literacy course can move self-assessed literacy substantially in a single semester. In the final survey, 145 respondents rated their agreement with ten statements about their understanding of AI twice: as they recalled their level before the course and as they judged it afterward. Every question showed a statistically significant gain, with the largest increases in the ability to list examples of AI, to discuss AI with an appropriate vocabulary, and to be literate about the technical components of AI. The authors take this as evidence that the course achieved its primary learning objective, and they use the participant feedback to identify what worked (varied speakers, concrete examples) and what did not (challenging readings, disconnected fundamentals).","pith_inferences":["A natural next experiment is to administer the same ten items as a true pre-test at the start of the course; comparing those baseline scores with the retrospective 'before' ratings would directly quantify response-shift bias.","The course's structure—short lectures by rotating experts plus weekly reflections—could transfer to workplace continuing-education or public-library settings, where AI literacy gaps are similar but credit and grading are absent.","If the reported gains reflect genuine learning, the ten-item retrospective instrument could serve as a lightweight evaluation tool for other institutions, but only after being validated against an objective knowledge measure."],"forward_implications":["The same 14-week lecture structure with interdisciplinary speakers can be re-deployed quickly at other institutions, since the paper shows it can be assembled in about six weeks.","A three-credit expansion of the course, using the same topic list with more interactive components, is already justified by the feedback and was offered in fall 2024.","The ten survey items provide a reusable, statistically significant outcome measure for evaluating future AI literacy courses.","The finding that lectures were rated easier than readings suggests future iterations should keep lecture-based fundamentals and replace or supplement technical readings with more accessible journalism."],"supporting_citations":[{"why":"Defines response-shift bias, the measurement pitfall the retrospective pre/post design is meant to avoid.","marker":"Howard and Dailey 1979"},{"why":"Provides the methodological rationale for retrospective pre-post surveys, the instrument behind the paper's central finding.","marker":"Geldhof et al. 2018"},{"why":"The AI100 report is the course's foundational reading and sets the scope of AI literacy the curriculum addresses.","marker":"Littman et al. 2022"},{"why":"The closest prior university AI literacy course for diverse students, which the paper positions as its baseline for comparison.","marker":"Kong, Cheung, and Zhang 2021"},{"why":"Supplies the qualitative coding approach used to identify themes in open-ended participant feedback.","marker":"Auerbach and Silverstein 2003"},{"why":"The coding manual applied to analyze free-text survey responses and derive the lessons learned.","marker":"Saldaña 2021"}],"fun_headline_variants":["All ten AI-literacy measures rise after 14-week course","One course lifts AI literacy on every survey question","145 learners report broad AI-knowledge gains in one course","AI literacy course yields up to 1.37-point gains for all"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's central claim rests on retrospective self-reports: participants rated their own literacy before and after the course in a single sitting, with no control group or objective knowledge test to confirm that self-assessed gains correspond to real learning.","fun_headline_variants_meta":{"raw":{"variants":["All ten AI-literacy measures rise after 14-week course","One course lifts AI literacy on every survey question","145 learners report broad AI-knowledge gains in one course","AI literacy course yields up to 1.37-point gains for all"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000188,"raw_usage":{"total_tokens":1306,"prompt_tokens":891,"completion_tokens":415,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":346}},"tokens_in":507,"tokens_out":415,"duration_ms":5309,"temperature":1.0,"reasoning_tokens":346,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:42:21.246733+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct the same ten-item instrument as a true pre-test at the start of the course and again at the end, with a no-course control group; if the course group shows no larger gain than the control, or if the true pre-test ratings already match the retrospective 'before' ratings, the reported improvement is largely a measurement artifact.","supporting_citations":[{"cited_title":"S.; and Dailey, P","cited_arxiv_id":null,"evidence_quote":"Defines response-shift bias, the measurement pitfall the retrospective pre/post design is meant to avoid."},{"cited_title":"J.; Warner, D","cited_arxiv_id":null,"evidence_quote":"Provides the methodological rationale for retrospective pre-post surveys, the instrument behind the paper's central finding."},{"cited_title":"Gathering Strength, Gathering Storms: The One Hundred Year Study on Artificial Intelligence (AI100) 2021 Study Panel Report","cited_arxiv_id":"2210.15767","evidence_quote":"The AI100 report is the course's foundational reading and sets the scope of AI literacy the curriculum addresses."},{"cited_title":"M.-Y.; and Zhang, G","cited_arxiv_id":null,"evidence_quote":"The closest prior university AI literacy course for diverse students, which the paper positions as its baseline for comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the qualitative coding approach used to identify themes in open-ended participant feedback."}],"review_version":1}