{"id":"e9122bec-a3b1-4cf8-b931-e16a8a462225","arxiv_id":"2505.03867","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An embedded AI copilot helped children brainstorm, debug, and create images in a Scratch-like environment, while children frequently rejected or adapted suggestions to preserve creative control.","lead":"Researchers built an AI chat assistant into a Scratch-like coding environment for children and watched 18 kids ages 7 to 12 use it. The children mainly used the assistant to brainstorm ideas, debug code, and generate images, and they often rejected or changed its suggestions to keep control of their projects.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AI-vs-researcher attribution is unresolved: the tool's author scaffolded every session and intervened after AI failures, with inconsistent counts (Section 3.2 vs 4.2), so 'the tool can effectively scaffold' is not yet demonstrated.","rationale":"The reader's weakest assumption correctly identifies the attribution problem as the most load-bearing issue. The paper is an exploratory qualitative study, so it does not need a controlled efficacy trial, and its codebook and quotes are internally coherent. However, the conclusion uses 'demonstrated,' which is stronger than the method supports. The researcher who built the tool was present in every session, prompted AI use, and provided direct help after AI failures; the amount of that help is reported inconsistently across Section 3.2 (~10 instances) and Section 4.2 (20 failures, 'two to three times per session'). The first author also transcribed and translated all non-English data and coded most transcripts, so the analysis path is not independent of the intervention. Because the central claim is exactly that the AI copilot can scaffold creative coding while preserving agency, this conflation is not a peripheral limitation; it changes what can be concluded. The concern is addressable by re-coding the existing video data to separate AI-only from researcher-assisted exchanges, so the correct verdict remains conditional rather than rejection. The reader's conditional verdict already reflects this, so no change is needed.","tokens_in":17731,"tokens_out":6041,"duration_ms":58044,"concrete_test":"Re-analyze the 20 recorded sessions by creating a timestamped log of every researcher utterance and every AI response, then label each coded instance (Code Support, Design Support, Conceptual Support, Child Agency, AI Failure) as AI-only or researcher-assisted. Recompute the code frequencies and the reported 70% success estimate for AI-only exchanges only. If most supporting quotes and agency examples occur in researcher-assisted exchanges, the conclusion must be revised from 'the tool can effectively scaffold' to 'an AI-plus-researcher system scaffolded, with the AI's independent contribution unmeasured.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the exploratory study 'demonstrated that such a tool can effectively scaffold the creative coding process' requires the observed scaffolding to be attributable to the AI copilot. The method in Section 3.2 makes the researcher part of the intervention: the researcher suggested AI use and offered direct assistance after two failed AI responses, in 'approximately 10 instances.' Section 4.2 then says the first author intervened for the 20 AI-failure instances, 'occurring two to three times per session on average,' which is numerically inconsistent with 10 total instances and implies a much larger human role; the 70% success estimate is also presented without a calculation basis. Section 3.5 adds that the first author conducted, transcribed, and translated all sessions and coded most transcripts. The evidence therefore supports an 'AI-plus-researcher' system, not the copilot alone, and even the child-agency examples include researcher push-back prompts ('What do you think will happen if we ask again?'). The qualitative themes remain plausible, but the attribution needed for the headline claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents Cognimates Scratch Copilot, an AI assistant embedded in a block-based Scratch-like environment, and reports an exploratory qualitative study with 18 children (ages 7–12) from 11 countries. Each session had a pre-AI coding phase, an AI-enhanced coding phase in which the researcher encouraged AI use and offered direct help after two failed AI attempts, and a semi-structured reflection interview. Thematic analysis of video transcripts produced three themes: AI-enhanced ideation and asset creation, contextual debugging and system navigation, and preservation of child agency. The paper proposes design guidelines and concludes that the tool can effectively scaffold youth creative coding by aiding ideation, debugging, asset creation, and platform navigation while preserving child agency.","tokens_in":17894,"tokens_out":4297,"duration_ms":38368,"significance":"The contribution is potentially valuable for the IDC/HCI community: it is one of the first empirical accounts of an LLM-based copilot embedded in a visual programming environment for children, and it provides concrete observational evidence of youth agency, over-reliance concerns, and failure-driven learning. Strengths include the explicit exploratory framing, a defined codebook with occurrence counts, illustrative quotes, an international sample, and attention to culturally responsive design. However, the paper's headline claim currently exceeds what the method can support, because the study conflates the AI tool with substantial researcher scaffolding and does not measure creative self-efficacy directly.","major_comments":[{"comment":"The attribution of the observed scaffolding to the AI tool alone is not established. Section 3.2 states that researcher scaffolding occurred in approximately 10 instances, while Section 4.2 reports 20 AI failure instances with the first author intervening 'two to three times per session on average'—a numerical inconsistency (for 18–20 sessions, 20 failures is roughly one per session, not two to three). The first author also conducted, transcribed, translated, and mostly coded all sessions (Section 3.5) and actively suggested AI use and clarified AI responses (Sections 3.2 and 4.2). The evidence therefore describes an AI-plus-researcher intervention, not the copilot alone. To support the conclusion's claim that the tool 'can effectively scaffold' the creative coding process, the authors must either separate AI-only from researcher-assisted interactions in the analysis or explicitly reframe the conclusion to describe the combined system.","section":"Sections 3.2, 4.2"},{"comment":"The estimate that 'the AI successfully answered about 70% of queries' is undefined and unverifiable. The paper does not define what counts as a query, what counts as success, or how the estimate was calculated across sessions. As this figure is used as a quantitative anchor for the AI's effectiveness, the authors should either provide a clear definition and calculation basis, or remove the estimate and restrict claims to the qualitative patterns.","section":"Section 4.2"},{"comment":"The paper repeatedly invokes creative self-efficacy, but no measure of creative self-efficacy appears in the method or codebook; the reported evidence consists of observed behaviors, interview statements, and parent emails. The abstract appropriately uses 'potential to enhance,' but the conclusion states the tool has 'the potential to empower youth, enhance their creative self-efficacy' and the introduction claims the study builds on prior work 'while promoting creative self-efficacy' without measuring it. These claims should be explicitly labeled as hypotheses or directions for future work unless a validated self-efficacy instrument is included.","section":"Sections 1, 5.2, 6"}],"minor_comments":[{"comment":"The text refers to 'the second and third author,' but the paper lists only two authors; this should be corrected to match the author list or to identify the specific coders by name.","section":"Section 3.5"},{"comment":"Participant J. is described as age 14 (New Zealand), outside the stated 7–12 age range; please verify and correct the age or the stated range.","section":"Section 4.1"},{"comment":"Reference [27] appears truncated and malformed; the citation should be completed.","section":"References"},{"comment":"The text reads 'In 9 of 18 cases, children explicitly declined an AI's suggestion,' which is ambiguous; if the intended meaning is '9 of 18 participants,' it should say so to match the later phrasing.","section":"Section 4.3"},{"comment":"Typo: 'a participatory design study the involved children' should be 'that involved children.'","section":"Section 2.2"},{"comment":"Possessive apostrophes should be consistent: 'childrens'' should be 'children's.'","section":"Section 3.4"}],"recommendation":"major_revision","confidential_remarks":"The heavy reliance on the authors' own prior work (references 13–17) is not disqualifying given the continuity of the research line, but the introduction's 'first tool' claim should be checked against any concurrent or prior systems. The main barrier to acceptance is the attribution gap between the AI-only claim and the AI-plus-researcher procedure, along with the undefined 70% success estimate and the unmeasured creative self-efficacy construct. These are fixable with reframing and additional reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: if you work on child-AI interaction or AI education, this is worth a read. It reports a real system, real sessions with 18 kids from 11 countries, and a codebook-grounded qualitative analysis of how children used the AI copilot. The agency finding — kids rejecting, adapting, or preempting AI suggestions — is the strongest and most interesting part. The design guidelines at the end are reasonable and emerge from the data.\n\nWhat's new: a GPT-4o/DALL-E-3 copilot embedded in a Scratch-like environment, with a system prompt that asks guiding questions rather than giving answers. That's a genuine application shift from adult copilots. The cross-cultural sample is unusually diverse for a study this size, and the quotes give a concrete sense of how 7–12-year-olds negotiate AI suggestions.\n\nNow the soft spots. The stress-test note is right: the attribution is not clean. Section 3.2 says researcher scaffolding occurred in roughly 10 instances, but Section 4.2 says the first author intervened for 20 AI-failure instances, 'two to three times per session on average.' Those numbers don't line up, and the first author ran every session, prompted kids to use the AI, and stepped in after two failed AI responses. So the evidence supports an AI-plus-researcher system, not the copilot alone. The conclusion's 'demonstrated that such a tool can effectively scaffold' is too strong for exploratory data with that confound.\n\nAlso: creative self-efficacy is claimed in the abstract and conclusion, but never measured directly. The 70% AI success rate is presented without a calculation basis. And the 'first tool' contribution is hard to square with the cited prior work, including the authors' own Scratch Copilot Evaluation [16] and Kazemitabaar's middle-school code generation studies. The unresolved overlap with [16] deserves a sentence in the revision.\n\nThe qualitative themes — ideation support, debugging help, and especially child agency — remain plausible and well-illustrated. The flaws are addressable, not fatal. With a revised framing that presents this as an exploratory study of an AI-plus-researcher scaffolding intervention, and with the count inconsistencies fixed, this would be a solid contribution.\n\nFor peer review: yes, send it out. It's a serious, readable HCI paper with honest limitations. I'd accept as a conditionally accepted paper after the authors fix the attribution language and the numerical inconsistency. I probably wouldn't cite it in my own work in the next year unless I was writing specifically about child-AI copilots, in which case the agency findings would be useful.\n\nReading group: maybe — good for a children-and-AI reading group discussion about method and attribution.","headline":"A useful exploratory study of an AI copilot for kids' block-based coding, but the headline claim overreaches: the researcher was part of the intervention, so the paper demonstrates an AI-plus-researcher system, not the copilot alone.","tokens_in":18427,"tokens_out":2109,"would_cite":false,"duration_ms":18875,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that an AI copilot embedded in a block-based coding environment can scaffold children's creative coding—ideation, debugging, asset creation, and platform navigation—without eroding their creative ownership.","keywords":["AI copilot","creative coding","children","Scratch","block-based programming","child agency","scaffolding","qualitative study"],"falsifier":"Run the same 40–50 minute creative coding protocol with no researcher present and no fallback help after failed AI responses, then compare ideation success, code completion, agency behaviors such as rejections and adaptations, and self-reported creative self-efficacy against the current sessions. If the observed scaffolding and agency behaviors disappear or drop sharply without the researcher, the claim that the copilot alone carries the effect is falsified.","tokens_in":17509,"feed_emoji":"🤖","tokens_out":7218,"duration_ms":68212,"temperature":0.7,"pith_summary":"Translating an imaginative idea into working code is a known hurdle for children learning to program. This paper argues that an AI assistant placed inside a block-based coding environment can share that burden—generating project ideas, suggesting code and debugging steps, creating images on request, and guiding platform navigation—while leaving the child as the creative decision-maker. The argument rests on an exploratory study of 18 children ages 7–12 from 11 countries using a purpose-built assistant called Cognimates Scratch Copilot. The authors report that children used the assistant as a first resort for help, took its suggestions when useful, and adapted or rejected them often enough to keep authorship and control. If the claim holds, AI copilots can be designed for children that scaffold rather than substitute, a useful result for creative computing education.","feed_headline":"AI copilot helps kids code without seizing control","feed_subtitle":"In a study of 18 children ages 7–12, the assistant aided ideation, debugging, and images while kids often said no.","key_machinery":"The carrying mechanism is the copilot's question-driven dialogue protocol embedded in the system prompt. The assistant is told to keep responses to a short child-friendly phrase, ask a guiding question before answering, and only give the answer if the same question comes more than twice; for code questions it gives one specific tip or one guiding question, and for image requests it calls a separate image-generation model. This protocol is what lets the tool scaffold without taking over: it forces the child to be the one who tries, evaluates, and decides, which the paper links to the observed agency behaviors. A second mechanism is the side-by-side interface—chat window and block canvas visible together—so prompts and generated images can feed directly into the project.","core_discovery":"The central claim is that a question-first AI copilot, not a one-shot answer machine, can support the full arc of a child's creative coding project. Designed as a chat pane beside a Scratch-like block canvas, the copilot uses a system prompt that instructs it to respond with a single short tip or a guiding question, to withhold the answer until the child has asked the same thing more than twice, and to offer ideation, code explanation, and image generation. In sessions with 18 children, the authors observed it supporting brainstorming (13 of 18 children asked for ideas), code and debugging help (46 coded instances), visual asset creation (33 instances), and platform navigation (12 instances), with children using it 3–12 times per session. Agency was not lost: 9 of 18 children explicitly rejected at least one suggestion, often describing themselves as the captain of the project, and approximately 30 percent of queries were unsuccessful—yet those failures became moments of prompt refinement and conceptual insight rather than dead ends. The paper concludes that such a tool can effectively scaffold creative coding while children maintain creative control.","pith_inferences":["The strongest untested extension is whether the same agency and learning outcomes survive without the researcher in the room: the sessions included the tool's builder, who suggested AI use and stepped in after two failed responses, so the measured effect is plausibly an AI-plus-human team rather than the copilot alone.","A natural next design is a context-aware copilot that can see the sprite, blocks, and screen state; children themselves requested this, and it would likely cut the observed approximately 30 percent failure rate from ambiguous queries.","Age differences hinted at in the data—younger children conversing socially, older children wanting more control and advanced features—could be tested with age-stratified studies and might lead to copilot personas that adapt their scaffolding style by developmental stage.","The same question-first architecture could be tested beyond Scratch-like blocks, such as in text-based introductory programming or creative tools outside coding, where the need to scaffold without eroding agency is similar."],"forward_implications":["Children in the 7–12 age range can treat an AI assistant as a resource to query, negotiate with, and override, rather than an authority to obey, which is a precondition for using copilots in creative learning settings.","Question-first system prompts are a transferable design choice: they can be adopted by other block-based coding tools to give help while preserving problem-solving practice.","Integrated image generation lowers the barrier to visual asset creation in children's coding projects, so designers should treat asset creation as a core copilot function rather than an add-on.","AI failures, when framed as part of the interaction, can produce teachable moments such as prompt refinement and debugging insight, so copilot designs should make errors visible and recoverable rather than hidden.","The proposed design guidelines—prioritize agency, balance support and challenge, allow customization and multimodal input—offer a concrete checklist for future youth AI coding tools."],"supporting_citations":[{"why":"This prior participatory design study supplied the design goals and system-prompt preferences, such as guiding questions and positive feedback, that define the copilot persona.","marker":"[15]"},{"why":"This defines the Scratch block-based environment that the copilot extends and the creative coding context studied.","marker":"[32]"},{"why":"This documents usability challenges novice programmers face when first using Scratch, the problem the copilot is meant to address.","marker":"[23]"},{"why":"This provides prior evidence on code generation tools' effects on middle schoolers' programming performance, the closest baseline the study extends toward younger children.","marker":"[24]"},{"why":"This supplies the metacognitive scaffolding rationale for asking guiding questions rather than giving direct answers.","marker":"[29]"},{"why":"This classification of creativity support tools grounds the claim that effective support must be integrated into ideation, implementation, and reflection.","marker":"[21]"},{"why":"This thematic analysis method was used to derive the three findings themes from session transcripts.","marker":"[8]"},{"why":"This creative self-efficacy measure informed the study's coding frame and the claimed outcome of enhanced creative confidence.","marker":"[50]"}],"fun_headline_variants":["AI copilot helps kids code, but kids keep control","Scratch copilot: kids lead, AI follows","Children steer AI assistant in coding study","AI coding aid for kids: scaffolding, not takeover"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes the effects it reports came from the copilot, even though the researcher who built the tool was present in every session, encouraged children to use it, and directly helped after two failed AI responses.","fun_headline_variants_meta":{"raw":{"variants":["AI copilot helps kids code, but kids keep control","Scratch copilot: kids lead, AI follows","Children steer AI assistant in coding study","AI coding aid for kids: scaffolding, not takeover"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000262,"raw_usage":{"total_tokens":1622,"prompt_tokens":996,"completion_tokens":626,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":564}},"tokens_in":612,"tokens_out":626,"duration_ms":6308,"temperature":1.0,"reasoning_tokens":564,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:44:50.475819+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 40–50 minute creative coding protocol with no researcher present and no fallback help after failed AI responses, then compare ideation success, code completion, agency behaviors such as rejections and adaptations, and self-reported creative self-efficacy against the current sessions. If the observed scaffolding and agency behaviors disappear or drop sharply without the researcher, the claim that the copilot alone carries the effect is falsified.","supporting_citations":[{"cited_title":"AI Friends: A Design Framework for AI-Powered Creative Programming for Youth","cited_arxiv_id":"2305.10412","evidence_quote":"This prior participatory design study supplied the design goals and system-prompt preferences, such as guiding questions and positive feedback, that define the copilot persona."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This documents usability challenges novice programmers face when first using Scratch, the problem the copilot is meant to address."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This supplies the metacognitive scaffolding rationale for asking guiding questions rather than giving direct answers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This classification of creativity support tools grounds the claim that effective support must be integrated into ideation, implementation, and reflection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This creative self-efficacy measure informed the study's coding frame and the claimed outcome of enhanced creative confidence."}],"review_version":1}