{"id":"724b5f48-c511-4a3d-9d8f-f2a84468214b","arxiv_id":"2412.11911","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Hour of Code AI activities for middle and high school students overwhelmingly emphasize perception and machine learning, neglect representation and reasoning, and offer limited hands-on engagement.","lead":"This study reviewed 47 Hour of Code AI and machine learning activities for middle and high school students, finding they mostly teach perception and machine learning while giving little attention to representation and reasoning. It matters because Hour of Code reaches over 1.8 billion activities, so these patterns shape what millions of youth first learn about AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'telling vs. doing' finding—a central part of the paper's claim—is asserted without reported counts, an explicit operational definition, or per-activity data, so this key result cannot be verified from the manuscript.","rationale":"The reader's verdict of CONDITIONAL is reasonable, and I do not dispute the overall direction of the paper. My stress-test surfaced a different point than the reader's stated weakest assumption. The reader worries that public activity content may not reflect actual classroom implementation, but the paper's claims are framed as content analysis—how activities engage with the five big ideas—and the authors explicitly limit their conclusions and call for observational follow-up. The topic-coverage percentages are internally consistent and grounded in the reported coding. By contrast, the 'telling > doing' component is presented as a key finding: it appears in the abstract, the findings, and the discussion, yet the manuscript reports no numbers for it. I looked through the Methods and Findings for definitions or counts of hands-on versus telling and found only illustrative examples. This is not a dispute with the qualitative conclusion; it is a reproducibility gap at the level of the paper's own central dependent variable. A concrete, small re-analysis of the 47 public activities would settle whether the claim lands. If the authors can provide the per-activity instructional-method codes, or the codebook and a second coder's application, the CONDITIONAL verdict can be satisfied; otherwise the 'limited hands-on engagement' finding remains an unsupported impression. The concern is about the argument's evidentiary support, not about the authors' integrity; no ad hominem is intended or implied.","tokens_in":10364,"tokens_out":4828,"duration_ms":46273,"concrete_test":"Make the instructional-approach coding explicit and public: define 'hands-on' (e.g., learner manipulates inputs, builds/retrains a model, writes/runs code, or experiments with data) versus 'telling' (e.g., concept conveyed only via video, slides, or text), have two independent coders apply this scheme to all 47 activities, and report per-activity codes together with the existing big-idea and ML codes. Provide counts of hands-on, telling, and both for the 37 activities, plus inter-rater agreement (e.g., Cohen's kappa). Then compare the counts with the Discussion's 'telling > doing' statement. If the counts show that a majority of activities include any hands-on component, the claim would need to be softened or replaced with a distribution; if the counts corroborate the qualitative assertion, the finding becomes reproducible.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim has two parts: topic coverage (perception and ML dominant, representation low) and an instructional-approach claim that activities are far more often 'telling' than 'doing.' The first part is backed by percentages (83.33%, 75%, 41.67%, 13.89%). The second part is not. In the Methods, the authors list instructional methods as an inductively coded category, but the manuscript never defines what counts as 'hands-on' versus 'telling,' nor the unit of coding. In Findings, the statement that ideas were often integrated 'through videos or teacher explanations without opportunities for students to engage in hands-on ways' is illustrated with examples, but no table, count, or figure legend quantifies how many of the 37 activities were hands-on, telling, or both. The Discussion's 'we noted how much more telling than doing was chosen' is phrased as a result, yet the supporting evidence is not reported. Because this is a content analysis of publicly available activities, classroom observation is not required for the instructional-approach claim; what is required is a systematic, reproducible coding of the activity materials themselves. The reader's concern about classroom implementation is valid for generalizing to learning outcomes, but the 'telling vs. doing' finding is about the activities as designed, so the more direct and load-bearing vulnerability is the missing data for this central dependent variable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a content analysis of Hour of Code (HoC) activities related to artificial intelligence and machine learning (AI/ML), focusing on beginner activities for middle and high school students. Using the \"five big ideas\" framework of Touretzky et al. (2019) and a companion framework for machine learning (Touretzky et al., 2023), the authors coded 47 AI-related activities identified from the HoC website in December 2023 and April 2024. They report that 37 activities engaged with at least one big idea, with perception (83.33%) and learning (75%) most common, followed by natural interaction (41.67%) and societal impact (41.67%), and representation and reasoning least common (13.89%). The paper also describes how ML is introduced (e.g., definitions, training data, algorithms, bias, training/test phases), discusses societal and environmental impact topics, and claims that instructional approaches were much more \"telling\" than \"doing\" (hands-on). The discussion offers design recommendations for future introductory AI/ML activities.","tokens_in":10568,"tokens_out":4388,"duration_ms":37779,"significance":"If the findings are accurate, this is a useful and timely empirical map of a widely used outreach program, contributing to the growing literature on K–12 AI literacy. The use of an external framework (Touretzky et al.'s five big ideas) rather than an author-defined scheme is a strength, and the authors are appropriately careful to describe the analysis as exploratory and descriptive rather than causal. The increased attention to societal impact compared with earlier HoC analyses is a notable and policy-relevant observation. However, the paper's central quantitative claims are currently undermined by an internal denominator inconsistency and by the absence of any reported data behind the \"telling vs. doing\" finding, which limits the verifiability of the main conclusions.","major_comments":[{"comment":"The percentages reported in this section appear to use a denominator of 36, not the stated 37 activities. For example, 30 activities addressing perception would be 81.08% of 37 but 83.33% of 36; similarly 27/36 = 75%, 15/36 = 41.67%, and 5/36 = 13.89%. The manuscript should either correct the number of activities (e.g., 36) or report the correct percentages for 37, because these percentages are the central descriptive result of the paper.","section":"Findings, 'How Did AI HoC Activities Address the Five Big Ideas of AI?'"},{"comment":"The claim that activities are much more \"telling\" than \"doing\" is stated in the Findings, abstract, and Discussion, but the manuscript never defines the categories \"hands-on\" and \"telling,\" does not specify the unit of coding (e.g., activity, segment, or topic), and provides no counts, table, or per-activity breakdown. As a result, a key finding cannot be verified from the reported evidence. The authors should add an operational definition of these categories, describe how they were applied, and report the number of activities (or idea-instances) coded as telling, hands-on, or both, ideally broken down by big idea.","section":"Analysis and Findings, 'telling vs. doing'"},{"comment":"The procedure for identifying AI/ML activities from the large HoC repository is underspecified. The text states that the authors \"reviewed the short description of each activity to select those mentioning AI/ML,\" but it does not describe the search or screening criteria, how the 557 (or 542) descriptions were processed, or the specific method by which nine unlabeled activities were found. This lack of detail limits reproducibility and makes it difficult to assess potential selection bias in the resulting 47-activity corpus.","section":"Data Collection"},{"comment":"The statement that \"nearly one-third\" of AI/ML-related HoC activities addressed societal impact contradicts the reported figure of 41.67% (15 of 36 activities). The text should be corrected to \"over two-fifths\" or the percentage should be revised.","section":"Discussion, 'The Why'"}],"minor_comments":[{"comment":"The title asks \"What Can Youth Learn ... in One Hour?\" but the study analyzes activity materials, not actual learning outcomes. The authors do acknowledge this limitation in the Discussion, but the title and abstract could be more precise (e.g., \"What Do Hour of Code Activities Aim to Teach...\").","section":"Title and Abstract"},{"comment":"The relationship among the numbers 47, 38, 28, 10, and 9 is confusing. The paper initially says 47 AI-related activities were selected, then later reports 38 labeled AI activities (28 with big ideas, 10 without) and 9 unlabeled activities with big ideas (28+9=37). Please clarify the arithmetic and explain why the total corpus is 47 while the analysis of big ideas is based on 37 (or 36).","section":"Findings, first paragraph"},{"comment":"The phrase \"Using the Touretzky and colleagues' (2023) five big ideas\" is a citation error: the five big ideas are from Touretzky et al. (2019), while the 2023 reference is the companion paper on machine learning. Please correct this citation.","section":"Discussion, first paragraph"},{"comment":"These figures use colors to indicate whether big ideas were incorporated into hands-on activities, telling activities, or both, but the figure captions do not explain the color legend. Since the text refers to these colors, the captions should include an explicit legend or refer to a common color key.","section":"Figure 3 and Figure 5"},{"comment":"The claim that \"all the HoC activities were individual in nature\" is made without presenting any coding or analysis that would support it. If this was part of the inductive coding, the relevant category and findings should be reported; otherwise, the statement should be softened or removed.","section":"Discussion, 'Offering Unplugged and Collaborative AI/ML Activities'"},{"comment":"The comparison between the 41.67% of AI-specific activities addressing societal impact and the less than 2% of all HoC activities in earlier years is not apples-to-apples, because the earlier figure applies to the general HoC catalogue, not to AI-related activities. The authors should note this difference in scope when interpreting the improvement.","section":"Discussion, 'The Why'"}],"recommendation":"major_revision","confidential_remarks":"The paper's topic is well within the scope of the journal and the empirical corpus is relevant. The core finding about perception/ML dominance is plausible and likely correct, but the manuscript currently has a few load-bearing reporting gaps (denominator inconsistency, missing operationalization of 'telling vs. doing', and underspecified selection process) that prevent the reader from verifying the central claims. These are all fixable with additional tables, definitions, and text revisions. I do not see a fundamental flaw that would require rejection; the authors should be encouraged to revise."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, clearly-scoped content analysis that gives the field its first systematic map of the AI/ML activities in Hour of Code. The finding that perception and learning dominate while representation and reasoning are nearly absent is well supported by the reported percentages, and it is useful for anyone designing or evaluating these outreach materials.\n\nThe paper's real contribution is the corpus itself: 47 AI-related activities identified through a transparent filtering process, coded against the established five big ideas and Touretzky et al.'s ML aspects, with additional inductive coding for societal-impact topics. The unusually high attention to societal impact, relative to earlier HoC analyses, is a genuinely interesting result. The recommendations—unplugged activities, attention to algorithms, better tools for model building—are pragmatic and follow from the data.\n\nThe soft spots are mostly about reporting. The 'telling vs. doing' claim is central to the discussion, but the manuscript never gives a numeric breakdown. Figure 3 and Figure 5 apparently encode the hands-on/telling distinction, but the prose should state how many of the 37 activities fell in each category and how the categories were operationally defined. Without that, the claim 'much more telling than doing' reads as an impression rather than a finding. That is the most load-bearing weakness, and it is fixable with a table.\n\nSecond, the coding method is collaborative consensus with no inter-rater reliability. That is defensible for an exploratory study, and the authors are transparent about it, but it means the analysis is hard to reproduce. A shared codebook and coding spreadsheet would help a lot. The related concern—that public activity content may not match what learners actually experience—is real but secondary; the authors acknowledge it and correctly limit their claims.\n\nMinor: the percentages for the five big ideas (83.33%, 75%, 41.67%, 13.89%) are consistent with 36 activities, but the text says 37. Someone needs to reconcile the denominator.\n\nOverall, the paper is honest and well within its scope. It deserves a serious referee and, after a small revision that quantifies the instructional-approach claims and shares coding artifacts, would be a useful addition to the CS-education literature.","headline":"Useful, clearly-scoped content analysis with credible topic-distribution findings, but the central 'telling vs. doing' claim needs explicit numbers and the coding artifacts should be shared.","tokens_in":11142,"tokens_out":3390,"would_cite":true,"duration_ms":30089,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Hour of Code AI activities cluster on perception, skip reasoning, and tell more than they do.","keywords":["Hour of Code","AI literacy","K-12 education","machine learning education","five big ideas of AI","content analysis","societal impact","computing education"],"falsifier":"Re-code the 47 activities with independent coders who do not discuss their judgments, and observe a sample of classroom implementations; if independent coding or classroom observation shows far more hands-on model building than the consensus coding captured, the central 'mostly telling' claim would be undercut.","tokens_in":10136,"feed_emoji":"🤖","tokens_out":8325,"duration_ms":63752,"temperature":0.7,"pith_summary":"The paper analyses all 47 Hour of Code activities that introduce artificial intelligence or machine learning to middle and high school students. It finds that the portfolio is lopsided: perception appears in 83% of the activities that engage with at least one big idea, learning in 75%, natural interaction and societal impact in about 42% each, and representation and reasoning in only 14%. Only two activities touch all five big ideas, and many activities labeled as AI contain no AI content at all. On pedagogy, the authors find far more 'telling' than 'doing': ideas are usually conveyed through videos and teacher guides rather than through learners building, training, or experimenting with models. The paper argues this matters because these short activities are many students' first and possibly only formal contact with AI.","feed_headline":"Hour of Code AI lessons shortchange reasoning and hands-on work","feed_subtitle":"Analysis of 47 activities: perception in 83%, representation in 14%, and societal impact on the rise.","key_machinery":"The analytical machinery is the five big ideas of AI—perception, learning, natural interaction, representation and reasoning, societal impact—used as a content taxonomy, together with a five-aspect breakdown of machine learning (defining ML, how learning algorithms work, the role of training data, data bias, and learning-versus-application phases). The authors apply this taxonomy deductively to the 47 activities, add inductive codes for societal-impact topics and instructional mode (telling vs. hands-on), and code collaboratively for unanimous consensus. The taxonomy is what turns a collection of tutorials into comparable measures of content coverage and pedagogical engagement.","core_discovery":"The paper's central claim is that the Hour of Code's AI and machine learning offerings, as a portfolio in 2023–2024, are narrow in content and shallow in hands-on engagement. Using the five big ideas of AI as a coding scheme, the authors report that of the 37 activities addressing at least one big idea, perception is most common (83.33%), followed by learning (75%), natural interaction (41.67%) and societal impact (41.67%), with representation and reasoning last (13.89%). A deeper look at machine learning shows that while 17 activities ascribe model behavior to training data, only nine address learning algorithms, and ten of the 27 ML-related activities never define machine learning. On the instructional side, the authors note that ideas are often integrated through videos or teacher explanations without opportunities for hands-on engagement, and that several activities labeled as AI-related by the platform do not engage with any big idea. The increase in attention to societal impact—15 of 37 activities—is described as a surprising and welcome shift compared with earlier Hour of Code offerings.","pith_inferences":["A testable extension of the 'telling vs. doing' finding is to measure learning outcomes: for example, whether students who train a model with Teachable Machine for five minutes can explain how training data shapes behavior better than students who only watch a video.","Because the authors coded public activity content rather than classroom implementations, their percentages may overstate actual engagement; observing real classrooms could push the balance even further toward 'telling.'","The finding that only two activities integrate all five big ideas suggests a design challenge worth addressing directly: a one-hour activity that touches all five ideas may need to be a structured game or unplugged simulation rather than a screen-based tutorial.","A follow-up design study could test whether activities that foreground the energy cost of training models—currently absent from the portfolio—change students' attitudes toward AI's environmental impact."],"forward_implications":["If the analysis holds, the Hour of Code's flagship AI offerings give millions of students a lopsided first view of AI: heavy on perception and data-centric machine learning, light on reasoning and representation.","Some activities labeled as AI by the platform teach no AI at all, which risks misleading students, educators, and parents about what AI is.","The jump from six AI activities in 2021 to 47 by April 2024 shows fast growth, but the concentration of content means scale has not yet brought breadth.","The rise in societal-impact coverage, from under 2% in earlier Hour of Code activities to more than 40% of AI activities, suggests critical perspectives are becoming part of introductory AI education.","The authors' recommendations imply that a single hour can be better used with unplugged activities, collaborative data work, and novice tools for actually building models."],"supporting_citations":[{"why":"Supplies the five big ideas of AI taxonomy that is the study's primary coding scheme.","marker":"(Touretzky et al. 2019)"},{"why":"Provides the five-aspect account of machine learning the authors use to code ML content.","marker":"(Touretzky, Gardner-McCune, and Seehorn 2023)"},{"why":"Earlier Hour of Code analysis establishing that under 2% of activities engaged critical issues, the baseline for the reported rise in societal impact.","marker":"(Morales-Navarro et al. 2021a)"},{"why":"Earlier analysis of Hour of Code engagement modes that frames the telling-versus-doing concern.","marker":"(Morales-Navarro, Kafai, and Gregory 2022a)"},{"why":"Introduces Teachable Machine, the hands-on tool the authors cite as a successful model for involving novices in training models.","marker":"(Carney et al. 2020)"},{"why":"Systematic review of Hour of Code research cited to note that little is known about outcomes beyond participation numbers.","marker":"(Yauney, Bartholomew, and Rich 2023)"}],"fun_headline_variants":["Hour of Code AI: perception dominates, reasoning scarce","One-hour AI lessons favor perception, shortchange reasoning","AI Hour of Code: hands-on rare, societal impact rising","Hour of Code AI neglects reasoning, limits hands-on"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that the public descriptions, videos, and tutorials of each activity, as read and coded by three researchers in consensus, faithfully represent what learners actually experience and can learn in the hour.","fun_headline_variants_meta":{"raw":{"variants":["Hour of Code AI: perception dominates, reasoning scarce","One-hour AI lessons favor perception, shortchange reasoning","AI Hour of Code: hands-on rare, societal impact rising","Hour of Code AI neglects reasoning, limits hands-on"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000503,"raw_usage":{"total_tokens":2437,"prompt_tokens":903,"completion_tokens":1534,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":519,"completion_tokens_details":{"reasoning_tokens":1468}},"tokens_in":519,"tokens_out":1534,"duration_ms":11319,"temperature":1.0,"reasoning_tokens":1468,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:26:57.846034+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-code the 47 activities with independent coders who do not discuss their judgments, and observe a sample of classroom implementations; if independent coding or classroom observation shows far more hands-on model building than the consensus coding captured, the central 'mostly telling' claim would be undercut.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the five big ideas of AI taxonomy that is the study's primary coding scheme."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces Teachable Machine, the hands-on tool the authors cite as a successful model for involving novices in training models."},{"cited_title":"Hour of Code","cited_arxiv_id":null,"evidence_quote":"Systematic review of Hour of Code research cited to note that little is known about outcomes beyond participation numbers."}],"review_version":1}