REVIEW 3 major objections 5 minor 14 references
A Systematic Mapping Study on Chatbots in Programming Education
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims that the published literature on chatbots for programming education is dominated by introductory, Python-focused tutors, with generative LLM-based interaction models now the leading design and advanced programming topics…
desk verdict A serviceable SMS whose trends are plausible, but the 34-vs-54 corpus mismatch undermines every reported proportion until fixed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the systematic mapping study protocol itself: a repeatable literature search and coding procedure. The authors define search terms over three concepts (chatbot, programming, student or learning), query ACM, Engineering Village, IEEE Xplore, and Scopus, apply inclusion and exclusion criteria, and then code the retained studies against five research subquestions covering chatbot types, programming content, languages, interaction models, and application contexts. The interaction-model coding uses a named three-way taxonomy attributed to Hien et al. (2018): pattern-based, retrieval-based, and generative, extended with hybrid cases. This machinery does the work of converting a dispersed set of primary studies into the claimed proportions and gaps. The pedagogical strategy categories are derived inductively from reading the studies, and the paper itself notes that most included systems lack explicit instructional grounding.
What would settle it
Rerun the stated search string in ACM, Engineering Village, IEEE Xplore, and Scopus using the paper's priority order and inclusion and exclusion criteria; if the selected set differs substantially from the reported 54 studies (or the 34 stated in the introduction), or if a random sample of 20 included studies is reassigned to different categories on re-coding, then the mapping's proportions and gaps do not reproduce.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is the synthesis itself: from 3,216 initial records the authors selected 54 studies (the introduction reports a smaller count of 34 from 2,497 analyzed records) and found that educational chatbots for programming are predominantly introductory, Python-centered tools. Most teach language-specific fundamentals such as variables, control structures, and syntax; only a handful address object-oriented programming, data structures, web development, or physical computing. The interaction model has migrated toward generative models powered by large language models, with pattern-based and retrieval-based systems forming the earlier layer and hybrid architectures emerging as a promising but rare combination. A secondary finding is that most reported chatbots lack explicit grounding in learning theories. The paper presents these as trends and gaps in the published record, not as experimental evidence about which chatbot design works better.
Load-bearing premise
The entire map depends on whether the search string, the four databases, and the inclusion and exclusion rules actually captured the relevant literature, and the paper itself gives conflicting counts of how many studies were included (54 in the abstract, 34 in the introduction), so every reported trend inherits that uncertainty.
Editorial extensions
If this is right
- New chatbot designs for programming education can be positioned against a known baseline: the default system in the literature is a Python-focused introductory tutor with generative interaction.
- Advanced topics such as data structures, object-oriented programming, web development, and physical computing are documented gaps; a designer covering these would be addressing a niche the current literature has not populated.
- The mapping's prevalence claims concern system descriptions, not measured learning outcomes, so outcome-focused evaluations of generative educational chatbots are a clear next step.
- Hybrid interaction architectures that combine pattern, retrieval, and generative components are described as promising but rare, which places hybrid design as an open research direction.
Reading between the lines
- If the corpus is representative, the concentration on Python suggests chatbot support follows the language of the most popular introductory course rather than being driven by pedagogical need; a testable extension is to compare chatbot coverage with enrollment-weighted language popularity in first-year programming courses.
- The near absence of data structures and object-oriented programming chatbots may partly reflect the limits of earlier rule-based systems; with generative models lowering those limits, a plausible next step is to watch whether LLM-based tutors migrate into advanced topics and whether their accuracy there is sufficient.
- The interaction-model axis may matter less for learning than the paper's pedagogical strategy categories; one way to test this is to compare learning outcomes across the five pedagogical categories while controlling for the interaction model.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports a Systematic Mapping Study (SMS) on the use of chatbots to support programming education in undergraduate courses. Following Kitchenham-style guidelines, the authors searched four digital libraries (ACM, Engineering Village, IEEE Xplore, Scopus) with a defined search string and applied inclusion/exclusion criteria, yielding a corpus of primary studies. They then categorize the studies along four research subquestions: chatbot types (pedagogical strategies), programming concepts addressed, programming languages taught, and interaction models. The central findings are that chatbots are predominantly used for introductory Python instruction, focus on fundamental programming concepts, increasingly rely on generative/LLM-based interaction models, and that advanced topics (data structures, OOP, web development) remain underexplored. The paper also identifies a lack of explicit pedagogical theory grounding in many designed systems and proposes future research directions.
Significance. If the underlying corpus and classifications are reliable, this mapping study provides a useful synthesis of a fast-growing area, offering a catalog of chatbots, a categorization of pedagogical approaches, and an identification of research gaps (e.g., scarcity of chatbots for advanced programming topics). The observed trends—Python dominance and the shift to LLM-based generative models—are consistent with the broader computing-education literature and the cited examples. The study follows standard SMS procedures with a documented protocol and transparent inclusion/exclusion criteria, which is a strength. However, the value of the synthesis depends critically on the correctness and completeness of the selected corpus. The manuscript's internal inconsistencies in corpus size and subquestion count, combined with the absence of a PRISMA figure and any inter-rater reliability measure, currently limit the confidence in the quantitative prevalence claims. These are fixable but load-bearing issues.
major comments (3)
- [Introduction vs. Abstract/§2.5] The paper reports two incompatible corpus sizes. The Introduction states that '2,497 retrieved studies' yielded '34 primary studies,' while the Abstract and §2.5 report 3,216 retrieved publications and 54 selected studies. Table 3 lists 54 studies, and all results in Section 3 are computed over 54 studies (e.g., 22 Python-tagged studies in §3.3 are 41% of 54 but 65% of 34). Because every prevalence claim in Sections 3–5 is a proportion of the selected corpus, the discrepancy makes the reported distributions unverifiable. The authors must reconcile these numbers and present a complete, auditable PRISMA flow diagram.
- [Abstract and Table 1] The Abstract promises 'five research subquestions,' but Table 1 lists only four (SQ1–SQ4) and the results sections address exactly four. No SQ5 is defined or analyzed. This is an internal inconsistency in the research protocol that needs to be resolved by either adding the missing subquestion or correcting the abstract.
- [§2.5–§3.4] The inductive taxonomy and coding decisions (pedagogical categories in §3.1, content categories in §3.2, language categories in §3.3, interaction models in §3.4) are presented without any inter-rater reliability or validation measure. Since the study's central claims are derived from the distribution of studies across these categories, the absence of a reliability assessment (e.g., Cohen's kappa on a sample) weakens the quantitative conclusions. A mapping study should either report such a measure or transparently discuss the consensus process in a way that allows readers to judge coding consistency.
minor comments (5)
- [§2.5] Figure 1, referenced as summarizing the selection process, is not present in the manuscript. The figure is essential for auditing the deduplication and selection stages, and its absence is a reproducibility gap that should be corrected.
- [§3.1] The text contains 'inumerous' which appears to be a typo for 'numerous.'
- [Table 3] The table header contains the typo 'Referece' instead of 'Reference.'
- [Throughout] The manuscript mixes English and Portuguese conventions, using 'e' instead of 'and' in several places (e.g., §2.2 'IEEE Xplore Digital Library (IEEE) e Scopus') and in many reference entries (e.g., 'Kitchenham e Charters 2007'). This should be standardized to English.
- [§3.2] In the OOP paragraph, the sentence 'These include S23, S43, and S52, often leveraging ChatGPT...' is grammatically incomplete and should be rephrased for clarity.
Circularity Check
No significant circularity: the SMS reports descriptive classifications, and no claim reduces to its own inputs or to a load-bearing self-citation.
full rationale
The paper is a systematic mapping study whose central claims are descriptive proportions computed over a corpus that it selected and coded. This is inductive synthesis, not a derivation that reduces to its inputs by construction. The pedagogical categories in §3.1 are explicitly said to be 'derived inductively through a systematic reading of the studies'; reporting that the classified studies fall into those categories is a summary of the coding, not a prediction fitted to an independent target. No parameter is fitted and then renamed as a prediction, and no 'uniqueness theorem' or similar result is imported from the authors' prior work. The reference list contains no load-bearing self-citation by the present author team; the cited interaction-model taxonomy (Hien et al., 2018) and the SMS guidelines (Kitchenham, Petersen, Kuhrmann) are external methodological sources. The serious problems in the paper are reporting inconsistencies: the Abstract and §2.5 state 3,216 retrieved and 54 selected studies, while the Introduction states 2,497 retrieved and 34 selected; Table 1 lists four research subquestions although the Abstract promises five; and the referenced Figure 1 is absent from the supplied text. These undermine the verifiability of the corpus proportions and belong in a correctness/completeness review, not a circularity analysis, because they do not show that any claimed trend is equivalent by construction to an input or to a self-citation. The central claims — Python predominance, focus on introductory concepts, and the shift toward generative models — are properties of the primary literature itself, not artifacts of the paper's own definitions; another researcher re-coding the same papers could disagree with some labels, which is an intercoder-reliability concern, not circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption The four chosen digital libraries (ACM, Engineering Village, IEEE Xplore, Scopus) and the stated search string retrieve the population of relevant studies on chatbots in programming education.
- domain assumption The inclusion and exclusion criteria (IC1-IC3, EC1-EC5) correctly separate studies about chatbots in undergraduate programming education from adjacent work.
- domain assumption The inductive six-category pedagogical taxonomy (content-based direct instruction, reflective/metacognitive tutoring, automated feedback, contextual/multimodal tutoring, autonomous self-study, domain-specific assistance) is a valid representation of the studies' designs.
- domain assumption The self-reported descriptions in the primary studies accurately reflect the chatbots actually deployed and evaluated.
Cite this review
Pith. "Pith review of A Systematic Mapping Study on Chatbots in Programming Education." pith.science (2026). https://pith.science/paper/4UI7FROR
@misc{pith2026250908857,
author = {Pith},
title = {Pith review of: A Systematic Mapping Study on Chatbots in Programming Education},
year = {2026},
howpublished = {\url{https://pith.science/paper/4UI7FROR}},
note = {Machine review of arXiv:2509.08857}
}
read the original abstract
Educational chatbots have gained prominence as support tools for teaching programming, particularly in introductory learning contexts. This paper presents a Systematic Mapping Study (SMS) that investigated how such agents have been developed and applied in programming education. From an initial set of 3,216 publications, 54 studies were selected and analyzed based on five research subquestions, addressing chatbot types, programming languages used, educational content covered, interaction models, and application contexts. The results reveal a predominance of chatbots designed for Python instruction, focusing on fundamental programming concepts, and employing a wide variety of pedagogical approaches and technological architectures. In addition to identifying trends and gaps in the literature, this study provides insights to inform the development of new educational tools for programming instruction.
Reference graph
Works this paper leans on
-
[1]
Introduction Programming poses challenges from both didactic and cognitive perspectives, hinder- ing its assimilation by students and its effective mediation by instructors [Robins 2019, Luxton-Reilly 2016]. As a result, introductory programming courses have historically shown high rates of failure [Alves et al. 2019] and dropout [Penney et al. 2023], und...
work page 2019
-
[2]
Research Method In this study, we conducted a Systematic Mapping Study (SMS) following the guidelines proposed by Kitchenham et al. (2007). Our goal is to investigate the chatbots described in the literature that support the teaching of programming in undergraduate courses. We detailed the procedures adopted in this study in the following subsections. 2.1...
work page 2007
-
[3]
Research subquestions and their motivations. Research Subquestions Motivation SQ1. Which chatbots have been proposed to support programming learning? SQ2. What programming concepts are addressed by educational chat- bots SQ3. Which programming lan- guages are employed in chatbot- mediated learning interactions? SQ4. What interaction strategies are used by...
work page 2010
-
[4]
The selected studies were published between 2003 and
Results Table 3 presents the complete list of selected studies. The selected studies were published between 2003 and
work page 2003
-
[5]
Final Considerations Our SMS investigated how chatbots have been developed and applied to support the teach- ing and learning of programming in undergraduate contexts. In response to the research question presented in Section 2.1, we found that chatbots have been predominantly em- ployed to support the learning of introductory programming concepts, especi...
work page 2024
-
[6]
ID Reference ID Reference ID Referece S01 [Coronado et al
Selected studies. ID Reference ID Reference ID Referece S01 [Coronado et al. 2018] S19 [Kumar et al. 2024] S37 [Wijaya e Purwarianti 2024] S02 [Verleger e Pembridge 2018] S20 [Gupta et al. 2025b] S38 [Vintila 2024] S03 [Lane e VanLehn 2003] S21 [Gupta et al. 2025a] S39 [Liu et al. 2024] S04 [Nguyen et al. 2022] S22 [Frankford et al. 2024] S40 [Lin 2022] S...
work page 2018
-
[40]
Modran, H., Ursutiu, D., Samoila, C., e Gherman Dolha˘scu, E.-C
Citeseer. Modran, H., Ursutiu, D., Samoila, C., e Gherman Dolha˘scu, E.-C. (2024). Developing a gpt chatbot model for students programming education. Em Advances in Intelligent Systems and Computing. Springer. Nguyen, H. D., Tran, T.-V., Pham, X.-T., Huynh, A. T., Pham, V. T., e Nguyen, D. (2022). Design intelligent educational chatbot for information ret...
work page 2024
-
[133]
Vadaparty, A., Geng, F., Smith IV, D
IEEE. Vadaparty, A., Geng, F., Smith IV, D. H., Benario, J. G., Zingaro, D., e Porter, L. (2025). Achievement goals in cs1-llm. Em Proceedings of the 27th Australasian Computing Education Conference, páginas 144–153. Van Merrienboer, J. J. e Sweller, J. (2005). Cognitive load theory and complex learning: Recent developments and future directions. Educatio...
work page 2025
Show all 14 references
-
[163]
Callejo, P., Alario-Hoyos, C., e Delgado-Kloos, C
Springer. Callejo, P., Alario-Hoyos, C., e Delgado-Kloos, C. (2024). Evaluating chatgpt impact on the programming learning outcomes of students in a big data course. International Journal of Engineering Education, 40(4):863–872. Carreira, G., Silva, L., Mendes, A. J., e Olivei...
2024 arXiv
-
[177]
e Pembridge, J
Verleger, M. e Pembridge, J. (2018). A pilot study integrating an ai-driven chatbot in an introductory programming course. Em 2018 IEEE frontiers in education conference (FIE), páginas 1–4. IEEE. Vintila, F. (2024). Avert (authorship verification and evaluation through respons...
2018
-
[327]
J.-K., Qiu, Z., Zhu, Y., Murnane, E
Ruan, S., Jiang, L., Xu, J., Tham, B. J.-K., Qiu, Z., Zhu, Y., Murnane, E. L., Brunskill, E., e Landay, J. A. (2019). Quizbot: A dialogue-based adaptive learning system for factual knowledge. Em Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pági...
2019 arXiv
-
[2018]
Initially limited to rule-based and narrowly scoped applications, contemporary chatbots have be- gun to incorporate more sophisticated approaches, expanding their roles beyond teaching basic syntax or commands. The presence of chatbots with multiple functionalities—such as emo...
2018
-
[2022]
In addition, the COVID-19 pandemic created an urgent need to restructure teaching and learning models, demanding accessible and interactive educational solutions
This increase may be associated with significant advancements in the field of Natural Language Processing (NLP), particularly following the introduction of the Transformer architecture in 2017, which enabled the development of more sophisticated chatbots with enhanced contextu...
2018
-
[2025]
A new study emerged in 2013, but once again, there was a hiatus until 2018, when we Figure
From the temporal perspective, the earliest relevant study was published in 2003, followed by a ten-year gap. A new study emerged in 2013, but once again, there was a hiatus until 2018, when we Figure
2003
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.