{"id":"50963f63-369a-4a94-98b0-67ecfd191bb0","arxiv_id":"2504.14769","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Three teens built a babyGPT screenplay generator in a five-day workshop, demonstrating feasibility of youth engagement in GLM design through data practices and ethical reasoning.","lead":"A case study of three teenagers shows they can meaningfully engage with generative AI by building a small 'babyGPT' trained on Marvel scripts, including data curation and ethical discussions about copyright. This suggests construction activities can move youth from using generative models to designing them.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Feasibility claim rests on treating dataset design and output evaluation as 'building'; actual model training was researcher-mediated, so the demonstrated scope is a data-pipeline construction plus evaluation, not the full training loop.","rationale":"The reader's weakest assumption is the same one I identified: the feasibility claim is weakened by the fact that actual training was performed by researchers, not by the youth. This is the most load-bearing concern because the paper's headline contribution is 'demonstrates the feasibility of engaging youth in building GLMs,' and the abstract says the team 'developed a model.' If the intended contribution is that youth can complete a model-building loop, the missing hands-on training is a direct gap. The concern is not that the observed activities are trivial—data curation, tokenization, prompt design, and output evaluation are real and well-documented—but that the scope of the claim is broader than the evidence. The paper is honest about the training arrangement in footnote 4 and about the group selection in Section 3.3, which is credit to the authors; the issue is calibration of the interpretation, not the quality of the case study. A single concrete check—whether any training episode appears in the video logs—would settle whether the youth engaged with training at all. If they did not, the abstract and discussion need to be tempered. Because the reader already issued a CONDITIONAL verdict with the same reasoning, my analysis does not change that verdict; I recommend UNCHANGED, meaning the reader's conditional acceptance remains appropriate.","tokens_in":9970,"tokens_out":2061,"duration_ms":21303,"concrete_test":"Re-examine the video logs and artifacts to determine whether any episode shows a participant initiating or observing a training run (e.g., accessing a terminal, seeing a loss curve, or discussing training iterations beyond filling the duration request form). If no such episode exists, the abstract and Section 5 should be revised to state that youth built the dataset and evaluated model outputs while training was performed by researchers; equivalently, the methods section should define 'building' to include researcher-mediated training.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that youth can build GLMs depends on what 'building' includes. Section 3.2, footnote 4, states that school administrators blocked terminal access and that all models were trained by a researcher after the workshop. The youth's only training-related decision was filling out a request form to choose training duration (5 minutes vs. 1 hour). The observed activities—curating and tokenizing screenplays, selecting prompts, and evaluating outputs—are genuine data practices, but they stop short of the training loop. There is no evidence in the paper that the youth saw loss curves, adjusted hyperparameters, or otherwise interacted with the training process; the Discussion itself lists 'explore validation and training loss, adjust weights, or finetune pre-trained models' as future work, conceding that these were absent. If 'building a GLM' requires hands-on training, then the case supports a narrower claim: youth can construct a dataset, specify a training request, and evaluate a researcher-trained model. The paper could repair this by explicitly defining building to include researcher-mediated training, or by softening the abstract's 'developed a model' and 'building GLMs' to 'constructed datasets and evaluated outputs of trained models.' The Section 3.3 choice of the group with the most complete data also makes this a best-case feasibility demonstration, which is acceptable for an exploratory case study but reinforces the need for scope-limiting language.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a five-day workshop in which 35 ninth-grade students designed and built small generative language models ('babyGPTs') using the nanoGPT framework. The authors focus on a single team of three teenagers who created a Marvel-screenplay generator, curating and tokenizing a dataset of roughly 80,000 tokens, selecting prompts, submitting a training request, and evaluating the outputs of the resulting model. Using video recordings, artifacts, and saved files, the authors construct a descriptive case study and analyze it with Olari and Romeike's inventory of AI/ML data practices. They report that the youth engaged in data collection, data quality control, data preparation, and performance evaluation, and that they grappled with ethical issues around copyright, authorship, and attribution. The paper's central claim is that this case study demonstrates the feasibility of engaging youth in building GLMs, positioning them as designers rather than merely users of generative AI.","tokens_in":10210,"tokens_out":4702,"duration_ms":42048,"significance":"If the central claim is accepted, this paper makes a useful contribution to child-computer interaction and AI/ML education by extending constructionist approaches from classification tasks to generative language models. The strength of the paper lies in its rich qualitative data: the authors provide detailed vignettes of youth decision-making about dataset composition, their reasoning about copyright and attribution, and their critical evaluations of model outputs. The use of an external data-practices framework (Olari and Romeike) and the availability of artifacts (datasets, model files, outputs) lend transparency to the analysis. The paper also honestly reports the institutional constraint that prevented youth from performing training themselves (footnote 4). However, the framing of what constitutes 'building' a GLM is broader than what was actually demonstrated, and this overreach affects the paper's central feasibility claim. The result is a plausible and valuable exploratory case study whose scope needs to be stated more precisely.","major_comments":[{"comment":"The central claim that the case study 'demonstrates the feasibility of engaging youth in building GLMs' overreaches the observed activities. Although the youth curated and tokenized a dataset, chose prompts, and evaluated outputs, all model training was performed by a researcher after the workshop because school administrators blocked terminal access; the youth's only training-related action was selecting a training duration on a request form. The Discussion itself defers 'explore validation and training loss, adjust weights, or finetune pre-trained models' to future work, which concedes that these parts of the training loop were absent. The abstract should be reworded to claim feasibility of youth constructing datasets, specifying training requests, and evaluating outputs of a researcher-trained model, or the paper should explicitly define 'building' to include researcher-mediated training.","section":"Abstract; Section 3.2, footnote 4; Section 5"},{"comment":"The case was chosen because it had the most complete data (filled-in worksheets, saved datasets, attendance). The paper presents this as a feasibility demonstration without acknowledging that this is a best-case selection; readers cannot tell whether the observed engagement is typical or even achievable by other groups. The Discussion should explicitly state that the claim is limited to a best-case demonstration and that future work is needed to test generalizability.","section":"Section 3.3"},{"comment":"The introduction states that participants engaged in 'implementing a solution' as one of the data practices, but the Findings describe the youth running a provided tokenization script and filling out a training request form; they did not implement the model or training pipeline. This is internally inconsistent and should be corrected either by removing 'implementing a solution' from the list of observed practices or by explaining how tokenization and training requests constitute implementation in this context.","section":"Section 1 and Section 4"}],"minor_comments":[{"comment":"The names 'Kaparthy' and 'Bathia' are misspelled; the references cite Andrej Karpathy and Aatish Bhatia.","section":"Section 3.2"},{"comment":"The text says 'see Figure 1' but the caption does not describe the output fragment; consider adding a caption that explains what is shown.","section":"Section 4, Figure 1"},{"comment":"The demographic description notes that all participants identified as White, but the paper does not discuss how this limits the generalizability of the findings; a sentence acknowledging this would be appropriate.","section":"Section 3.1"},{"comment":"The title and abstract refer to 'babyGPTs' (plural), but the case study follows one team building one model; consider clarifying whether the workshop had multiple groups and that the case is a single exemplar.","section":"Title and Abstract"},{"comment":"The sentence 'Our findings suggest that constructing GLMs should be an integral part of efforts to foster AI/ML literacies' is stronger than a single-case study supports; consider softening to 'may' or 'could'.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written and useful exploratory case study, but the authors need to tighten the scope of their feasibility claim. The main issue is not the quality of the observation but the gap between 'building a GLM' and what the youth actually did; this is fixable with careful wording. I would also encourage the authors to add a limitations paragraph covering the best-case selection and the lack of demographic diversity, even though neither is fatal for an exploratory study. The paper is potentially acceptable after revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. This is the first case study I know of where youth curate a corpus, tokenize it, and evaluate outputs of a small GPT they helped design, rather than just classifying or prompting. The novelty is modest but real: it moves constructionist AI literacy from classification to generative models. The paper is honest, clearly written, and the video/artifact evidence is used carefully. The authors anchor the coding in Olari and Romeike's data practices framework and use nanoGPT, so the analysis has an external backbone.\n\nThe soft spots are real but not fatal. The strongest is the 'building' language. Footnote 4 and the methods say the youth never trained the models; researchers ran nanoGPT overnight because school administrators blocked terminal access. The youth's training decision was choosing 5 minutes vs. 1 hour on a request form. So the demonstrated scope is dataset construction plus prompt design plus output evaluation. That is legitimately valuable, but the abstract's 'developed a model' and 'building GLMs' overstate it. The Discussion even lists exploring loss curves and adjusting weights as future work, conceding the gap. Rephrasing 'building' to include researcher-mediated training, or saying 'constructed datasets and evaluated outputs of trained models,' would fix it.\n\nAlso, the case was chosen for most complete data, and it is a single team of three White teenagers. That makes this a best-case feasibility demonstration, not evidence about typical youth. The authors acknowledge the exploratory design, so it is a limitation, not a hidden flaw. I would have liked inter-rater reliability on the coding or a shared dataset, but for a nine-page IDC paper this is within normal practice.\n\nVerdict: conditional accept, but closer to accept than reject. The central observation—that teens can engage meaningfully with data practices and ethical questions when given a dataset-to-evaluation pipeline—holds up. It is the scope of the claim that needs trimming. This paper deserves a serious peer review; a careful reviewer can push for the language fix without asking for a different study.\n\nWould I cite it? Probably not in my own work, but I'd bring it to a reading group on AI literacy. It is a solid, small step forward.","headline":"A genuinely novel but carefully scoped case study of teens building a dataset and evaluating a small GPT; the 'building' claim overreaches, but the paper earns a serious read.","tokens_in":10729,"tokens_out":2139,"would_cite":false,"duration_ms":19157,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A case study shows three teenagers successfully building a small generative language model, not just using one.","keywords":["GPT","language models","youth","LLMs","computational empowerment","data practices","machine learning","artificial intelligence"],"falsifier":"A direct test would run the same workshop with the training step left to the youth themselves (for example, on unlocked computers with a simplified training interface) and observe whether they can complete it; a finding that they cannot, or a replication with a randomly selected group showing no spontaneous data-quality or ethical reasoning, would narrow or undercut the feasibility claim as stated.","tokens_in":9747,"feed_emoji":"🎬","tokens_out":8091,"duration_ms":66265,"temperature":0.7,"pith_summary":"This paper reports a five-day workshop in which 35 ninth graders built very small generative language models, which the authors call 'babyGPTs', and then focuses on one three-person team that created a Marvel screenplay generator. The central claim is that the case demonstrates the feasibility of engaging youth as designers of GLMs, not merely as users: the three teenagers collected and curated screenplay scripts, tokenized them into an 80,050-token dataset, wrote prompt 'seeds', chose training durations, and then critiqued the model's outputs for missing structure and narrative coherence. Along the way they discussed whether it is acceptable to train on copyrighted screenplays and who deserves credit for generated text. The authors argue that this construction activity made the process of building a GLM transparent and allowed the youth to engage with data practices and ethical considerations concretely. If the feasibility claim holds, it suggests a path for AI literacy education that has young people building models rather than only interacting with them.","feed_headline":"Three teens built a babyGPT, not just used one","feed_subtitle":"A case study shows how 14-year-olds curated a screenplay dataset and reasoned about copyright while training their own model.","key_machinery":"The central object is the 'babyGPT' construction activity: a small generative language model that youth build by choosing a text domain, assembling a curated dataset of roughly 75,000 to 300,000 tokens, tokenizing it with a provided script, designing prompts as generation 'seeds', and submitting a training request specifying how long the model should train. The paper's analysis then codes video recordings and artifacts against an inventory of AI/ML data practices (cited as [38]) and a construction process framework (cited as [26]), which together supply the categories that turn the observed classroom activity into evidence that the youth engaged in data practices and ethical reasoning. The training itself is carried out by a researcher after the workshop using a lightweight GPT training framework (cited as [30]), so the machinery that carries the argument is the designed pipeline plus the analytic coding, not the model training itself.","core_discovery":"On its own terms, the paper's discovery is that three 14- to 15-year-olds, given a scripted construction pipeline, can carry out the human parts of building a small generative language model. The team of Dillpickle, Optimus, and Cyclops selected Marvel screenplays, debated which films to include and exclude on quality grounds, ran a tokenization script, monitored the resulting token counts, requested a 5-minute and a 1-hour training run, and then evaluated the trained model's outputs, observing that the text lacked screenplay structure and a coherent narrative arc. When asked about using copyrighted scripts, they reasoned about copyright, attribution, and authorship, with one member imagining selling the model's script to Marvel and reflecting that the AI, not a human writer, produced it. The authors frame this as evidence that youth can engage in ML data practices—collecting data, controlling data quality, preparing data, and evaluating performance—and that constructing GLMs is a feasible route to AI/ML literacy.","pith_inferences":["The paper's own evidence supports a narrower claim than the abstract's 'building GLMs': because the teenagers never ran the training, the demonstrated capability is building the data pipeline and evaluating outputs, not completing the full model-building loop. A stricter feasibility test would put the training step in the youth's hands.","The case was chosen for having the most complete data, so the demonstration likely represents a best case; whether typical groups engage this deeply is untested.","A testable extension would compare AI-literacy gains between youth who build babyGPTs and youth who only use commercial chatbots, measuring their later ability to explain how LLMs work or critique model outputs.","The youth's ethical reasoning, such as the scenario of selling a model-generated script to Marvel, suggests that building with copyrighted data can push adolescents toward nuanced thinking about derivative authorship; a future design could deliberately scaffold that into a fuller fair-use lesson."],"forward_implications":["If the feasibility claim holds, youth AI education has a working template: kids can meaningfully design small generative models through data curation, prompt design, and output evaluation even when the compute-heavy training step is handled by adults.","Construction activities of this kind can shift youth's stance toward generative AI from passive acceptance of outputs to critical examination—here, recognizing missing screenplay structure and narratively incoherent text.","Ethical questions of copyright and attribution arise naturally during dataset construction, which may make them more concrete and memorable than in abstract lessons.","The study motivates building easier-to-use tools for novices to train, validate, and fine-tune their own models, since the only barrier to hands-on training in this case was blocked terminal access at school.","It also indicates that youth's data practices are iterative and non-linear, which has design implications for how workshop time and scaffolding are structured."],"supporting_citations":[{"why":"The lightweight GPT training framework used to train the babyGPT models.","marker":"[30]"},{"why":"The babyGPT demonstration project that inspired the workshop's design and gave youth an example of small GPTs.","marker":"[5]"},{"why":"The inventory of AI/ML data practices used to code the video recordings.","marker":"[38]"},{"why":"The construction process framework used to organize the analysis of youth activity.","marker":"[26]"},{"why":"The computational empowerment framework that motivates engaging youth in construction activities.","marker":"[13]"},{"why":"The descriptive case study methodology underpinning the single-group design.","marker":"[46]"}],"fun_headline_variants":["Three teens train a babyGPT on Marvel scripts","Teens build babyGPT, shaping data and ethics","Case study: 14-year-olds construct their own GLM","Youth as AI builders: babyGPT case study"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that selecting, curating, and tokenizing a dataset, writing prompts, and evaluating outputs counts as 'building a GLM', even though the model training itself was performed by a researcher after the workshop rather than by the teenagers.","fun_headline_variants_meta":{"raw":{"variants":["Three teens train a babyGPT on Marvel scripts","Teens build babyGPT, shaping data and ethics","Case study: 14-year-olds construct their own GLM","Youth as AI builders: babyGPT case study"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0003,"raw_usage":{"total_tokens":1699,"prompt_tokens":877,"completion_tokens":822,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":771}},"tokens_in":493,"tokens_out":822,"duration_ms":7543,"temperature":1.0,"reasoning_tokens":771,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:40:38.074081+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would run the same workshop with the training step left to the youth themselves (for example, on unlocked computers with a simplified training interface) and observe whether they can complete it; a finding that they cannot, or a replication with a randomly selected group showing no spontaneous data-quality or ethical reasoning, would narrow or undercut the feasibility claim as stated.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The babyGPT demonstration project that inspired the workshop's design and gave youth an example of small GPTs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The descriptive case study methodology underpinning the single-group design."}],"review_version":1}