{"id":"ade7b610-f9bf-43b0-824a-6a3e5ffe7483","arxiv_id":"2502.09799","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A co-design study with K-12 teachers yields design guidelines for LLM tools that support project-based learning, prioritizing teacher creativity, agency, and ethical integration.","lead":"This paper describes how 41 K-12 teachers helped co-design ideas for LLM tools to support project-based learning, through interviews, workshops, and wireframe feedback. It proposes design guidelines for such tools, emphasizing teacher agency, privacy, and practical classroom constraints.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-selected, mostly STEM/PBL-enthusiastic sample is the load-bearing weakness; the K-12-wide guidelines rest on unrepresentative teacher input.","rationale":"The reader identified the weakest assumption as the generalizability of design guidelines from a self-selected, PBL-enthusiastic sample. I agree: this is indeed the most load-bearing assumption because the paper's contribution—a set of design guidelines for K-12 PBL—is only as credible as the representativeness of the teachers whose needs and values shaped those guidelines. The paper is methodologically transparent, reports interrater reliability, and honestly states its limitations, so there is no internal inconsistency or overreach relative to its scoped claim. The absence of a working prototype and the lack of classroom deployment are appropriate limitations for a design-guideline contribution. The proposed concrete test, recruiting a non-self-selected and more diverse sample and repeating the core co-design activities, would directly test whether the guidelines hold beyond the original sample. Because the paper already acknowledges this limitation and because the contribution is exploratory rather than evaluative, the reader's ACCEPT verdict remains appropriate; the concern does not require a change in verdict.","tokens_in":28936,"tokens_out":6507,"duration_ms":67282,"concrete_test":"Recruit a purposive sample of teachers who are not self-selected: for example, from school-district rosters rather than opt-in lists, balanced for non-STEM subjects, elementary grades, and low PBL/GenAI comfort. Repeat the wireframe feedback and one co-design workshop from Study 1. If the themes and design recommendations change materially, the original guidelines are sample-specific; if the same recommendations recur with no new load-bearing themes, generalizability is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the proposed design guidelines are grounded in K-12 teachers' needs—depends on the 41 teachers being sufficiently representative of K-12 educators. Study 1 (n=11) recruited self-selected teachers with high PBL comfort (mean 4.6/5); Study 2 (n=30) were STEM teachers admitted to MIT's SEPT program, a network likely predisposed to educational technology. Both samples skew toward PBL enthusiasm and GenAI interest, and the paper's Limitations section (Sec. 6) explicitly concedes the self-selected sample 'potentially biasing the findings toward a more positive outlook.' If the guidelines were generated by teachers who are already convinced of PBL and GenAI value, they may not transfer to teachers in traditional, resource-constrained, or non-STEM settings—the very contexts where LLM tool adoption is contested. This is a limitation, not a fatal flaw, because the paper scopes its contribution to design considerations rather than efficacy, but it is the weakest load-bearing assumption in the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a two-study co-design process with 41 K-12 teachers to understand the challenges of project-based learning (PBL) and to derive design guidelines for LLM-based tools that support PBL. Study 1 involved 11 expert PBL teachers in semi-structured interviews, two co-design workshops, storyboarding, and wireframe feedback; Study 2 involved 30 STEM teachers with varying PBL experience who reviewed the same wireframes. The authors identify three support areas—curriculum, assessment, and progress tracking—and propose eight design recommendations structured around the Buck Institute's Gold Standard PBL framework. The paper claims to be the first co-design study with K-12 teachers specifically targeting LLM tools for PBL.","tokens_in":29228,"tokens_out":7116,"duration_ms":71462,"significance":"The contribution is timely and valuable: it opens a design space that has largely been explored with students in higher education, and it grounds tool requirements in teachers' own accounts of PBL practice. The qualitative analysis is systematic and transparent: themes are defined in Table 3, interrater agreement is high (Cohen's kappa 0.906 on 40% of the data), and the path from participant quotes to themes to wireframes to design recommendations is generally traceable. The paper also deserves credit for clearly labeling the wireframes as a starting point rather than a validated product and for candidly acknowledging self-selection, the U.S.-only context, and the absence of racial demographic data. The skeptical concern about sample representativeness is real, but it lands as a scoping caveat rather than a correctness error: the paper's central contribution is a set of design considerations grounded in a particular group of teachers, and the limitations section already acknowledges the main threats to generalizability. The result, if adopted by the community, would provide an evidence-based starting point for teacher-centered LLM tool design in PBL settings.","major_comments":[],"minor_comments":[{"comment":"The manuscript consistently describes the sample as 'interdisciplinary K-12 teachers,' but 30 of the 41 participants were STEM teachers recruited through the MIT SEPT program, and the wireframe feedback from both studies is pooled to shape the guidelines. The limitations section acknowledges self-selection and U.S. scope but not the disciplinary skew. Please add a sentence noting that the cross-disciplinary recommendations rest primarily on the 11 Study-1 participants, and soften the abstract's 'interdisciplinary' phrasing accordingly.","section":"3.2.1 and 6"},{"comment":"The abstract and Section 1.1 describe 'iterative design of wireframes' and 'iterative feedback,' but the methods describe a single wireframe construction followed by one review round in Study 1 and one review round in Study 2. If the wireframes were revised between these rounds, please say so explicitly; otherwise replace 'iterative' with language such as 'feedback-based refinement' or describe the two-stage review process.","section":"1 and 3.1.2"},{"comment":"There is a duplicated word in the sentence 'Unlike problem-based learning, which promotes promotes deductive reasoning'; 'promotes' should appear once.","section":"2.1"},{"comment":"The definition for the 'specific project examples' theme contains a garbled fragment: 'Include details about the project sp we know what is, not just in reference to other parts pedagogy.' Please repair this sentence so that the coding definition is clear to readers and future coders.","section":"Table 3"},{"comment":"A Cohen's kappa of 0.906 is described as 'substantial interrater reliability'; under the commonly used Landis and Koch convention cited elsewhere in the paper, this value is conventionally called 'almost perfect.' Consider aligning the wording with the cited scale.","section":"3.1.3"}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid qualitative HCI contribution with a clear co-design process and honest limitations. The main risk is that the K-12-wide framing slightly overreaches the STEM-heavy, self-selected sample, but this can be fixed with targeted wording changes rather than new data collection. I see no integrity or scope-fit concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a well-run co-design study and a useful contribution to the HCI-education subfield. The genuinely new piece is the focus on K-12 teachers specifically—earlier co-design work on GenAI in PBL has mostly been with students in higher ed. The authors do a careful job: two studies (11 expert PBL teachers, then 30 STEM teachers from the MIT SEPT network), transparent recruitment, thematic analysis with a reported Cohen's kappa of 0.906, and clear mapping from interview themes through workshops to wireframes to design recommendations. The guidelines are concrete and practical—single-point rubrics as default, local data processing for IEP privacy, teacher-directed scaffolding, choice boards for differentiation. These are the kind of recommendations that ed-tech designers can actually use, and they are grounded in what teachers said rather than in speculation. The paper also explicitly scopes itself as design considerations, not efficacy claims, and the limitations section is honest about the small sample, the self-selection, and the lack of classroom deployment.\n\nThe soft spots are real but not load-bearing. The biggest concern is the sample: Study 1 recruited self-selected teachers with high PBL comfort (average 4.6/5), and Study 2 was all STEM teachers connected to MIT's SEPT program—a network that is likely predisposed to ed-tech and GenAI. So the guidelines are built from teachers who are, on the whole, already convinced of PBL and GenAI value. The authors concede this in their limitations but still frame the recommendations broadly as being \"for PBL\" in K-12. That framing is a bit broader than the evidence supports, especially for non-STEM, resource-constrained, or GenAI-skeptical contexts. Minor soft spots: the wireframes (CAIL) are not a working product, so the guidelines are validated only against teacher feedback on mockups; there is no student voice in the co-design; and the paper's claim of being \"first\" is plausible but depends on how the scope is drawn. None of this undermines the core contribution. The methods are reproducible, the analysis is honest, and the recommendations are clearly derived from the data.\n\nI'd send this to peer review without hesitation, and I'd expect it to be accepted after revisions that soften the generalizability claims and more clearly distinguish the SEPT sample from a general K-12 population. Worth a serious referee.","headline":"A solid, honest co-design study that gives ed-tech designers a practical, teacher-grounded set of guidelines for LLM tools in K-12 PBL, with the main caveat being the self-selected, PBL-friendly sample.","tokens_in":773,"tokens_out":797,"would_cite":true,"duration_ms":29084,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Co-design with 41 teachers yields eight guidelines for LLM tools that augment PBL teachers, not replace them.","keywords":["generative AI","large language models","project-based learning","co-design","K-12 teachers","design guidelines","assessment","progress tracking"],"falsifier":"A replication of the same co-design protocol with a sample that includes teachers who are skeptical or resistant to both PBL and generative AI, or with a demographically broader and non-self-selected group; if their expressed needs contradict the eight recommendations (for example, if they demand fully automated grading or reject LLM input into rubric creation), the generalizability of the guidelines collapses. Equally, a working implementation of the CAIL wireframes could be tested for whether teacher agency and workload actually improve in a controlled classroom trial.","tokens_in":28709,"feed_emoji":"🎓","tokens_out":3692,"duration_ms":31709,"temperature":0.7,"pith_summary":"The paper argues that large language models can ease the practical burdens of project-based learning (PBL) for K-12 teachers, but only if the tools are designed with teachers, not for them. Through interviews, two co-design workshops, and iterative wireframe feedback with 41 interdisciplinary U.S. teachers, it identifies the core PBL pain points—project design, assessment, and progress tracking—and translates them into a set of design guidelines for LLM tools. The central message is that teachers want LLMs to automate routine administrative work and enrich personalized learning while leaving pedagogical judgment, final grading, and creative direction in human hands. The paper positions this as the first co-design study with K-12 teachers specifically aimed at incorporating LLMs into PBL pedagogy.","feed_headline":"Eight guidelines for AI that helps, not replaces, PBL teachers","feed_subtitle":"K-12 teachers co-designed the rules for LLM tools that ease project-based learning's workload without taking over.","key_machinery":"The load-bearing mechanism is the multi-stage co-design process itself: semi-structured interviews with expert PBL teachers, two collaborative workshops using divergent and convergent design thinking (brainstorming, storyboards, and prototype worksheets), and iterative walkthroughs of wireframes for a teacher-facing system the paper calls CAIL (Collaborative AI for Learning). These wireframes—covering curriculum supports, assessment supports, and progress tracking—are the concrete artifacts that translate teacher values into design requirements. The paper then maps the resulting requirements onto the Buck Institute's Gold Standard PBL: Project Based Teaching Practices framework, which organizes the final eight recommendations.","core_discovery":"The central claim is that a teacher-driven co-design process yields a concrete, grounded specification for what LLM tools should do in project-based classrooms: augment teacher creativity in project ideation, scaffold lesson planning against standards, generate differentiated and equitable rubrics (defaulting to single-point rubrics), track student progress at individual, group, and class levels, and support—but never replace—teacher grading and feedback. Teachers in the study consistently advocated for tools that support their professional growth and augment their existing roles, while flagging privacy, equity, and over-reliance as key risks. From this, the paper proposes eight design recommendations mapped onto the Buck Institute's Gold Standard PBL teaching practices, covering curriculum support, assessment support, and progress tracking.","pith_inferences":["If these guidelines hold, a natural next step is co-designing the student-facing side of the same tools; the authors note this only as future work, but the guidelines imply the teacher dashboard's effectiveness depends on student-facing mirrors of progress data.","The recommendation to keep IEP data local suggests a testable extension: compare teacher trust and adoption between a local-processing version and a cloud-LLM version of the same differentiation feature.","The findings imply that LLM tools could inadvertently embed a specific pedagogical philosophy; designers who adopt single-point rubrics and choice boards are also adopting PBL values, and that alignment should be made explicit.","A concrete extension would be a longitudinal deployment study measuring whether the time saved by the tool actually shifts toward the creative and fulfilling aspects of teaching, which is the benefit the teachers said they wanted."],"forward_implications":["LLM tools for PBL should treat teacher input as mandatory for core design decisions and offer optional customization features, preserving teacher agency.","Rubric generation should default to single-point rubrics and pair LLM suggestions with required teacher feedback rather than automated grading.","Progress tracking should be visible to students individually and anonymously at class level, while protecting against administrator misuse of teacher performance data.","Secure data handling for differentiation must keep IEP and other sensitive student data local or inside a school network.","Lesson-planning supports embedded in existing standards templates can lower the barrier for novice PBL teachers while reducing experienced teachers' planning load."],"supporting_citations":[{"why":"Supplies the review of PBL implementation challenges that the interview protocol and research objectives build on.","marker":"[32]"},{"why":"Provides the Gold Standard PBL teaching practices framework onto which the final design recommendations are mapped.","marker":"[96]"},{"why":"Provides the divergent and convergent workshop guidelines that structure the co-design sessions.","marker":"[91]"},{"why":"Defines co-design of innovations with teachers, the methodology the study claims to be the first to apply to K-12 LLM-PBL design.","marker":"[108]"},{"why":"Represents the prior co-design study with students in project-based learning that this paper explicitly distinguishes itself from.","marker":"[147]"},{"why":"Provides evidence that involving teachers in co-design increases ownership and adoption, motivating the study's approach.","marker":"[13]"}],"fun_headline_variants":["Eight guidelines for AI that augments PBL teachers","Teachers help design LLM tools for project-based learning","LLM tools for PBL: co-designed with teachers, not replacing them","How teachers want AI to support project-based learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings rest on the assumption that 41 self-selected U.S. teachers—most already enthusiastic about PBL and GenAI—voice needs and values representative enough of the broader K-12 teaching population to ground general design guidelines; the paper itself flags this in its limitations.","fun_headline_variants_meta":{"raw":{"variants":["Eight guidelines for AI that augments PBL teachers","Teachers help design LLM tools for project-based learning","LLM tools for PBL: co-designed with teachers, not replacing them","How teachers want AI to support project-based learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000818,"raw_usage":{"total_tokens":3538,"prompt_tokens":857,"completion_tokens":2681,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":2614}},"tokens_in":473,"tokens_out":2681,"duration_ms":18257,"temperature":1.0,"reasoning_tokens":2614,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T20:25:36.859789+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication of the same co-design protocol with a sample that includes teachers who are skeptical or resistant to both PBL and generative AI, or with a demographically broader and non-self-selected group; if their expressed needs contradict the eight recommendations (for example, if they demand fully automated grading or reject LLM input into rubric creation), the generalizability of the guidelines collapses. Equally, a working implementation of the CAIL wireframes could be tested for whether teacher agency and workload actually improve in a controlled classroom trial.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Gold Standard PBL teaching practices framework onto which the final design recommendations are mapped."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the divergent and convergent workshop guidelines that structure the co-design sessions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines co-design of innovations with teachers, the methodology the study claims to be the first to apply to K-12 LLM-PBL design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents the prior co-design study with students in project-based learning that this paper explicitly distinguishes itself from."}],"review_version":1}