REVIEW 3 major objections 6 minor 25 references
LLMs to Support K-12 Teachers in Culturally Relevant Pedagogy: An AI Literacy Example
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read In a four-teacher pilot, an LLM tool called CulturAIEd raised every participant's self-assessed confidence in culturally responsive AI-literacy lesson design to 5 out of 5.
desk verdict A well-scoped pilot of a genuinely new tool, but the abstract's causal claim outruns the evidence; worth a serious referee with revisions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is CulturAIEd, a lightweight LLM-powered activity-adaptation tool whose three components work together: student demographics entered by the teacher, a culturally responsive teaching (CRT) checklist that structures two tiers of modification (basic and advanced), and rubric-based scoring with formative feedback plus a just-in-time chatbot coach. The checklist is the load-bearing scaffold: it defines what counts as moving from tokenistic cultural references to substantive integration, and the LLM uses it along with the demographic context to generate concrete examples, analogies, and revisions. The mechanism is therefore not raw text generation but a structured loop that ties cultural content to the teacher's own classroom and scores the result against the rubric.
What would settle it
Run a larger version of the same study with a control group that revises the same activities using only the CRT checklist, then compare pre/post confidence surveys and have blinded experts rate the adapted lessons. The claim fails if post-tool confidence does not exceed the control, or if expert ratings show no real improvement in cultural responsiveness.
Extended reading notes
Core claim
The central discovery is that CulturAIEd improved teachers' confidence in culturally responsive lesson design in one sitting with only a handful of teachers. All four participants rated themselves 5/5 after using the tool on tasks such as adapting a prompt-engineering activity and an AI-hallucination activity, where they had earlier expressed uncertainty, limited time, and fear of making shallow cultural assumptions. Teachers credited the tool's straightforward demographic input, its concrete content suggestions, and its rubric-based feedback for turning what felt like a large, vague effort into specific revisions. The paper is careful to call these outcomes promising trends rather than conclusive evidence of efficacy.
Load-bearing premise
The central claim rests on four teachers' own post-session confidence ratings and interview statements, with no control condition and no independent check of whether the adapted lessons were genuinely culturally responsive.
Editorial extensions
If this is right
- If the pilot's self-reports hold, schools can offer teachers a low-cost, immediate path to culturally responsive versions of AI literacy lessons that currently exist only in generic form.
- The basic-versus-advanced modification tiers give teachers a concrete way to see the difference between surface-level cultural references and meaningful CRP, which addresses the authenticity problem teachers described in their own schools.
- Embedding rubric feedback and coaching inside the planning workflow lets CRP professional development happen during lesson preparation rather than as separate training sessions.
- Because teachers described the demographic-input step as simple, the same design could generalize to other emerging subjects whose curricula lack culturally tailored resources.
- The authors expect larger controlled studies to determine whether the observed confidence gains translate into lasting changes in teaching practice and student outcomes.
Reading between the lines
- A test the paper does not run: independent CRP experts blind-rate the original and adapted activities to check whether self-rated confidence corresponds to actual gains in lesson quality.
- The design does not separate the LLM's content generation from the checklist's scaffolding; a control condition that provides only the checklist, or only an unstructured chatbot, would isolate the active ingredient.
- If the active ingredient is the checklist-plus-demographics prompt rather than the polished tool, the minimal useful intervention might be a reusable prompt template teachers can run with any general-purpose LLM.
- The study measures teacher self-efficacy, not student experience; connecting the tool to engagement, belonging, or learning outcomes is the evidentiary step the paper leaves for future work.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents CulturAIEd, an LLM-powered tool intended to help K-12 teachers adapt AI literacy activities through Culturally Relevant Pedagogy (CRP). The authors report an exploratory pilot with four teachers: participants completed a pre-survey, three adaptation phases (unaided, CRT-checklist-guided, and CulturAIEd-guided), a post-survey, and semi-structured interviews. The reported findings are increased self-reported confidence in identifying and making culturally responsive modifications, high perceived efficiency, and qualitative themes around demographics integration, rubric-based feedback, and concrete examples. The authors frame the study as a preliminary exploration and explicitly state in Section 5 that the results are 'promising trends rather than conclusive evidence,' calling for larger controlled studies in future work.
Significance. As a proof-of-concept, the paper makes a useful contribution by describing a concrete tool design that integrates an LLM, a CRT checklist, and student demographic input for a domain where culturally responsive AI literacy resources are scarce. The pilot includes honest reporting of limitations, including the small sample, the lack of a control condition, and the acknowledged risk that LLM outputs may reinforce stereotypes or routinize CRP. If the self-reported confidence gains were corroborated by independent artifact evaluation and a controlled design, the tool could offer a scalable, low-effort pathway for teachers to begin integrating CRP. However, the present evidence is exclusively self-report and the abstract's causal wording overstates what the design can support.
major comments (3)
- [Abstract; §3.2; §4.3] The abstract's claim that 'CulturAIEd enhanced teachers' confidence' is not warranted by the study design. In §3.2, participants first completed an unaided adaptation (Phase I), then a checklist-based adaptation (Phase II), and only then used CulturAIEd (Phase III), so the post-survey shift from 3-5 to all 5 reported in §4.3 is confounded with practice effects, scaffolding, and demand characteristics. There is no control condition, and Section 5 itself concedes that the results are 'promising trends rather than conclusive evidence.' The abstract and results should be revised to describe a self-reported confidence increase in a single-session exploratory pilot, not a causal effect of the tool.
- [§3.2; §4.3] The outcome measure is exclusively self-reported confidence and satisfaction; no independent evaluation of the adapted activities is reported. RQ1 asks about teachers' 'ability to design culturally responsive AI literacy activities,' yet no artifacts from Phase I through Phase III are scored against the CRT checklist or rated by external reviewers. The current evidence cannot distinguish improved capacity from increased familiarity or social desirability. I recommend adding at least a small artifact analysis, such as blind rating of pre/post adaptations, or explicitly limiting the paper's claims to perceived confidence.
- [§3.1; §5] The discussion acknowledges that LLM outputs may contain cultural biases or stereotypes, but the manuscript provides no audit of the content generated from the demographic inputs in this study. Because demographic-driven content generation is the tool's central feature, the absence of even illustrative CulturAIEd outputs makes it impossible to assess whether the generated cultural content was authentic or tokenistic. I suggest including representative output examples with a brief check against the CRT checklist, or explicitly stating that authenticity was not assessed and treating it as an open risk.
minor comments (6)
- [§3.1; §3.2] Several typographical and spacing issues need correction, including the stray space in 'T ool Design' and the missing spaces in 'InthisIRB' and '90-120 minute.'
- [Table 1] The 'Post-Survey Results' row mixes Likert-scale outcomes with interview quotes; the exact survey item wording and response scale should be provided, and '3/4 strongly agreed' should be tied to the specific item.
- [§2] The claim that AI literacy curricula lack cultural contextualization could be strengthened by citing specific existing curricula or standards beyond [9], [11], [19], and [20].
- [Abstract] The clause 'which could result in high implementation efficiency' is grammatically incomplete; it should be connected to the preceding clause or rewritten.
- [References] Some references have incomplete bibliographic details, such as [4] being a URL-only resource and [23] providing only a DOI; full entries would improve reproducibility.
- [Figure 2] The demo screenshot in Figure 2 should be checked for legibility, and its annotations should be explained in a descriptive caption.
Circularity Check
No circularity found: the paper is an empirical pilot whose central claim rests on self-report data, not on a derivation chain or a fit that reproduces its own inputs.
full rationale
This paper does not present a derivation, fitted model, or uniqueness argument, so the standard circularity patterns do not apply. The central claim is that CulturAIEd 'enhanced teachers' confidence' based on pre/post self-report surveys and interviews (Section 4.3). The outcome variable is teacher confidence, which is measured independently of the tool's internal rubric and checklist; no equation or construction makes the outcome equal to the tool's inputs. The CRT checklist from reference [4] is used both to generate suggestions and to provide rubric-based feedback, but the study's evaluation is not based on that rubric scoring the teachers' output; it is based on self-reported confidence and qualitative interview quotes. This creates a shared-framework dependence in the intervention, but it is not a circular reduction. The paper's own Section 5 explicitly cautions that the improvements are 'promising trends rather than conclusive evidence' and calls for control comparisons, which is a validity limitation rather than a circularity. The only self-citations are references [19] and [25], used to support background claims about AI literacy resource gaps and K-12 AI literacy initiatives; these are not load-bearing for the pilot's findings. There is no fitted parameter renamed as a prediction, no self-citation chain forcing the conclusion, and no ansatz imported via citation that substitutes for independent evidence. Therefore the appropriate score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Culturally Relevant Pedagogy as defined by Ladson-Billings and the CRT checklist are valid and appropriate frameworks for AI literacy instruction.
- domain assumption Teachers' self-reported confidence from pre/post surveys is a meaningful proxy for CRP implementation capability.
- domain assumption An LLM conditioned on student demographic characteristics can produce culturally relevant content without harmful stereotyping.
Cite this review
Pith. "Pith review of LLMs to Support K-12 Teachers in Culturally Relevant Pedagogy: An AI Literacy Example." pith.science (2026). https://pith.science/paper/U54KF4ZP
@misc{pith2026250508083,
author = {Pith},
title = {Pith review of: LLMs to Support K-12 Teachers in Culturally Relevant Pedagogy: An AI Literacy Example},
year = {2026},
howpublished = {\url{https://pith.science/paper/U54KF4ZP}},
note = {Machine review of arXiv:2505.08083}
}
read the original abstract
Culturally Relevant Pedagogy (CRP) is vital in K-12 education, yet teachers struggle to implement CRP into practice due to time, training, and resource gaps. This study explores how Large Language Models (LLMs) can address these barriers by introducing CulturAIEd, an LLM tool that assists teachers in adapting AI literacy curricula to students' cultural contexts. Through an exploratory pilot with four K-12 teachers, we examined CulturAIEd's impact on CRP integration. Results showed CulturAIEd enhanced teachers' confidence in identifying opportunities for cultural responsiveness in learning activities and making culturally responsive modifications to existing activities. They valued CulturAIEd's streamlined integration of student demographic information, immediate actionable feedback, which could result in high implementation efficiency. This exploration of teacher-AI collaboration highlights how LLM can help teachers include CRP components into their instructional practices efficiently, especially in global priorities for future-ready education, such as AI literacy.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
Prolific, https://www.prolific.com/
-
[3]
Streamlit, https://streamlit.io/
-
[4]
CulturallyResponsiveTeachingChecklist.https://reimaginingmigration.org/resource- items/cultural-responsive-teaching-checklist/ (Jan 2019)
work page 2019
-
[5]
https://doi.org/10.3102/0034654315582066
Aronson, B., Laughter, J.: The Theory and Practice of Culturally Relevant Ed- ucation: A Synthesis of Research Across Content Areas86(1), 163–206 (2016). https://doi.org/10.3102/0034654315582066
-
[6]
Research in Social Sciences and Technology9(1), 329–350 (2024)
Baytak, A.: The content analysis of the lesson plans created by chatgpt and google gemini. Research in Social Sciences and Technology9(1), 329–350 (2024)
work page 2024
-
[7]
Qualitative Research in Psychology3(2), 77–101 (Jan 2006)
Braun, V., Clarke, V.: Using thematic analysis in psychology. Qualitative Research in Psychology3(2), 77–101 (Jan 2006). https://doi.org/10.1191/1478088706qp063oa
-
[8]
https://doi.org/10.1177/2158244016660744 8 J
Byrd, C.M.: Does Culturally Relevant Teaching Work? An Examination from Stu- dent Perspectives6(3) (2016). https://doi.org/10.1177/2158244016660744 8 J. Wang et al
Show all 25 references
-
[9]
https://doi.org/10.1186/s40594-023-00418-7
Casal-Otero, L., Catala, A., Fernández-Morante, C., Taboada, M., Cebreiro, B., Barro, S.: AI Literacy in K-12: A Systematic Literature Review10. https://doi.org/10.1186/s40594-023-00418-7
-
[10]
Dee, T., Penner, E.: The Causal Effects of Cultural Relevance: Evidence from an Ethnic Studies Curriculum54(1), 127 (2017)
2017
-
[11]
In: Proceedings of the 27th ACM Conference on on Innovation and Technology in Computer Science Education Vol
Druga, S., Otero, N., Ko, A.J.: The Landscape of Teaching Resources for AI Education. In: Proceedings of the 27th ACM Conference on on Innovation and Technology in Computer Science Education Vol. 1. pp. 96–102. ACM (2022). https://doi.org/10.1145/3502718.3524782
2022
-
[12]
https://doi.org/10.1007/s13218-021-00737-3
Eguchi, A., Okada, H., Muto, Y.: Contextualizing AI Education for K-12 Stu- dents to Enhance Their Learning of AI Literacy Through Culturally Responsive Approaches35(2), 153–161 (2021). https://doi.org/10.1007/s13218-021-00737-3
2021 doi
-
[13]
Gay,G.:CulturallyResponsiveTeaching:Theory,Research,andPractice.Teachers College Press (2000)
2000
-
[14]
Hammond, Z., Jackson, Y.: Culturally Responsive Teaching and the Brain (2015)
2015
-
[15]
https://doi.org/10.1128/jmbe.v21i1.2097
Johnson, A., Elliott, S.: Culturally Relevant Pedagogy: A Model To Guide Cultural Transformation in STEM Departments21(1), 21.1.35 (2020). https://doi.org/10.1128/jmbe.v21i1.2097
2020 doi
-
[16]
In: Olney, A.M., Chounta, I.A., Liu, Z., Santos, O.C., Bittencourt, I.I
Laak, K.J., Aru, J.: Generative ai in k-12: Opportunities for learning and utility for teachers. In: Olney, A.M., Chounta, I.A., Liu, Z., Santos, O.C., Bittencourt, I.I. (eds.) Artificial Intelligence in Education. Posters and Late Breaking Results, Workshops and Tutorials, In...
2024
-
[17]
https://doi.org/10.3102/00028312032003465
Ladson-Billings, G.: Toward a Theory of Culturally Relevant Pedagogy32(3), 465– 491 (1995). https://doi.org/10.3102/00028312032003465
1995 doi
-
[18]
https://doi.org/10.1353/hsj.0.0038
Leonard, J., Napp, C., Adeleke, S.: The Complexities of Culturally Relevant Ped- agogy: A Case Study of Two Secondary Mathematics Teachers and Their ESOL Students93(1), 3–22 (2009). https://doi.org/10.1353/hsj.0.0038
2009 doi
-
[19]
fromunseenneedstoclassroomso- lutions
Li,H.,Xiao,R.,Nieu,H.,Tseng,Y.J.,Liao,G.:“fromunseenneedstoclassroomso- lutions”: Exploring ai literacy challenges & opportunities with project-based learn- ing toolkit in k-12 education. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 39, pp. 29145–291...
2025
-
[20]
In: Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems
Long, D., Magerko, B.: What is AI Literacy? Competencies and Design Considera- tions. In: Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. pp. 1–16. ACM (2020). https://doi.org/10.1145/3313831.3376727
2020
-
[21]
International Journal of Educational Leadership Preparation11(1), n1 (2016)
Mette, I.M., Nieuwenhuizen, L., Hvidston, D.J.: Teachers’ perceptions of culturally responsive pedagogy and the impact on leadership preparation: Lessons for future reform efforts. International Journal of Educational Leadership Preparation11(1), n1 (2016)
2016
-
[22]
00336882251313701
Moorhouse, B.L., Ho, T.Y., Wu, C., Wan, Y.: Pre-service Language Teachers’ Task-specific Large Language Model Prompting Practices p. 00336882251313701. https://doi.org/10.1177/00336882251313701
-
[23]
https://doi.org/10.1186/ISRCTN13420346
Roy, P., Staunton, R., Poet, H.: ChatGPT in lesson preparation - A Teacher Choices Trial. https://doi.org/10.1186/ISRCTN13420346
-
[24]
Srate Journal27(1), 22–30 (2018)
Samuels, A.J.: Exploring culturally responsive pedagogy: Teachers’ perspectives on fostering equitable and inclusive classrooms. Srate Journal27(1), 22–30 (2018)
2018
-
[25]
In: European Conference on Technology Enhanced Learning
Tseng, Y.J., Yadav, G., Hou, X., Wu, M., Chou, Y.S., Chen, C.C., Wu, C.C., Chen, S.G., Lin, Y.J., Liao, G., et al.: Activeai: The effectiveness of an interactive tutor- ing system in developing k-12 ai literacy. In: European Conference on Technology Enhanced Learning. pp. 452–...
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.