REVIEW 3 major objections 5 minor 20 references
This paper surveys Arabic chatbots in education and finds only 10 unique systems as of August 2024, mostly retrieval-based and evaluated by human feedback.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A survey of 10 educational Arabic chatbots finds most are retrieval-based, use Modern Standard Arabic, and rely on human feedback rather than automatic metrics.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection A four-page update of two 2022 reviews that is fine as a quick overview but does not make its central scarcity claim auditable: no search protocol, no enumeration of the 8 inherited systems, and the key table is missing. the 3 major comments →
Arabic Chatbot Technologies in Education: An Overview
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that, as of August 2024, only 10 unique educational Arabic chatbots have been described in the literature. These break down into 7 retrieval-based systems, 2 framework-based systems, and 1 generative system. Almost all use Modern Standard Arabic, with only one supporting Classical Arabic and one supporting a Saudi dialect. Furthermore, most evaluations rely on human satisfaction measures rather than standard automatic metrics such as accuracy, F1-score, precision, or BLEU. The paper derives this inventory by combining the results of two 2022 reviews with two additional recent bots, and it interprets the outcome as evidence that educational Arabic chatbots remain
What carries the argument
The central object is the survey inventory itself: a table of the 10 identified chatbots, classified along three axes—adopted approach (retrieval, framework, generation), language variety (Classical Arabic, Modern Standard Arabic, dialect), and evaluation metric type (human-based or automatic). The argument works by aggregating two prior reviews, updating them with a 2022–2024 search, and using these dimensions to expose patterns of scarcity and immaturity.
Load-bearing premise
The count of ten systems is only as good as the two 2022 reviews' coverage plus the paper's own choice of just two newer bots; if either review missed systems, or if more post-2022 bots exist, the scarcity conclusion needs revision.
What would settle it
Conduct a fresh systematic search for educational Arabic chatbots from 2022 to the present, including non-academic and deployed systems; finding even a handful of additional bots—especially a second generative one—would weaken the scarcity claim and the maturity assessment.
If this is right
- If the count of 10 is right, there is a clear opening to apply generative and deep-learning techniques to Arabic educational chatbots, which are currently dominated by retrieval-based methods.
- The field would benefit from a unified benchmark with standard automatic metrics, since current reliance on human feedback makes systems hard to compare objectively.
- Classical Arabic and dialectal chatbots are almost nonexistent, pointing to the need for new large Arabic corpora to support them.
- Recent advances such as GPT and BERT have not yet translated into Arabic educational chatbot development, suggesting a delay in technology transfer.
- Researchers who want to contribute could focus on replacing subjective human satisfaction surveys with more rigorous and reproducible evaluation.
Where Pith is reading between the lines
- The scarcity may be partly an artifact of under-documentation: many working educational Arabic chatbots deployed in institutions or industry may never appear in academic literature, so the true count could be higher than 10.
- Since framework-based bots already exist, upgrading them to generative models could provide a relatively fast path to closing the maturity gap, a step the paper does not explicitly advocate.
- A testable extension would be to build a dialectal Arabic educational chatbot and compare student engagement against an MSA-only version, probing the paper's observation that dialects are rarely supported.
- Automatic metrics like BLEU and F1 may not capture pedagogical quality, so a hybrid benchmark combining task completion, learning gains, and user experience would be a stronger standard than either approach alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This short book chapter surveys Arabic educational chatbots. It builds on two 2022 reviews, adds two systems published between 2022 and August 2024, and reports a total of 10 unique educational Arabic chatbots. The authors classify these systems by approach (7 retrieval-based, 2 framework-based, 1 generative), language variety (mostly Modern Standard Arabic), and evaluation metrics (mostly human feedback). They conclude that educational Arabic chatbots are still scarce and mostly immature and recommend wider use of deep learning, automated metrics, and dialect/Classical Arabic corpora.
Significance. If substantiated, the survey would fill a small but real gap: it is one of the few works explicitly focused on Arabic chatbots in education, and it draws attention to the evaluation-metric problem in this area. The authors deserve credit for explicitly noting that paper counts may overstate system counts and for distinguishing system-level from publication-level analysis. However, the current significance is limited because the central inventory of 10 systems is not auditable: the supporting table is absent, the new-system search is undocumented, and the deduplication of the two source reviews is not shown. The conclusions about scarcity and immaturity rest entirely on this inventory, so the paper is not yet a reliable reference point.
major comments (3)
- [§3, Table 1] The central count of 10 systems is not verifiable from the submitted text. The sentence 'The following Table 1 shows a summary of all the bots found' is followed by a table header with no rows or data. Without a row-by-row inventory listing each system, its reference, approach, language variety, evaluation metric, and educational context, the claims that 7 are retrieval-based, 2 are framework-based, 1 is generative, and that MSA dominates cannot be checked. This is load-bearing because the paper's main finding ('scarce and mostly immature') is exactly this set of counts. Please provide the complete table and relate each row to the references.
- [§2–§3] The paper states that a new search was needed from 2022 to August 2024, but it never describes the search procedure. No databases, query terms, inclusion/exclusion criteria, language restrictions, or screening steps are given. Consequently, the assertion that only two systems ([4] and [8]) were added in that period is unsupported. In particular, the paper cannot rule out recent LLM-based educational Arabic tutors. A reproducible methods subsection (or at minimum an explicit search strategy) is required before the scarcity conclusion can be accepted.
- [§3, References [10]–[20]] The aggregation of the two reviews into '8 unique educational Arabic chatbots' is opaque. The reference list contains at least four clusters of papers that may describe the same underlying system: Abdullah ([10]–[11]), ArabChat ([12]–[14]), Aljameed/LANAI ([15]–[16]), and SIAAAC/SEG-COVID ([19]–[20]). The manuscript needs an explicit mapping from references to unique systems and a stated deduplication rule (e.g., same institution/authors/name). Without this, the '8 unique' count—and hence the total 10—cannot be independently reconstructed.
minor comments (5)
- [Abstract] Grammar: 'We were able to identified' should be 'We were able to identify'.
- [§1, last paragraph of p. 11 and p. 12] The citation for the three categories of Arabic is [4], which in the reference list is a specific chatbot paper (Alazzam et al., 2023). A general linguistic reference would be more appropriate for this taxonomy.
- [§3] 'Automated-based metrics (Accuracy, F1-score, precision, and BLUE)' contains a typo: the metric is BLEU, not BLUE.
- [§2] The novelty claim 'To the best of our knowledge, this is the first survey that focuses on Educational Arabic chatbots' should be softened or substantiated by a brief comparison with the two cited reviews and other educational-chatbot surveys. As written, it is stronger than the evidence provided.
- [References] Reference formatting is inconsistent: some entries lack volume, issue, or page ranges, and some have trailing periods in DOIs. Please harmonize with the journal style.
Circularity Check
No circularity: survey aggregates external literature; missing Table 1 and undocumented search are transparency limitations, not circular derivation.
full rationale
This paper is a literature review, not a derivation or modeling exercise. Its central claims—that only 10 unique educational Arabic chatbots were identified as of August 2024, that 7 are retrieval-based, 2 use frameworks, 1 is generative, that nearly all use MSA, and that most evaluations rely on human feedback—are summaries of externally published systems found in two prior reviews ([5] and [9]) plus two additional papers ([4] and [8]). No equation, model, or fitting procedure is present, and no parameter is fitted to a subset of data and then renamed as a prediction. The authors do not cite their own prior work, so there is no self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation. The 'first survey focusing on educational Arabic chatbots' remark is a novelty assertion, not a load-bearing circular premise. The manuscript itself points out limitations: the total count depends on the completeness of the two 2022 reviews, the selection of only [4] and [8] as new additions is not justified by a documented search strategy, and Table 1—which would contain the itemized evidence—is absent from the provided text. These are serious reproducibility and transparency concerns, but they are not circularity: the paper's conclusions are not equivalent to its inputs by construction, nor are they forced by self-reference. Therefore, the appropriate circularity score is 0.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption The two 2022 reviews ([5] and [9]) provide a comprehensive base of Arabic chatbots.
- domain assumption The categorization of each bot (technique, language variety, metrics) as reported in the original papers is accurate.
- ad hoc to paper Only two relevant educational Arabic chatbots were published between 2022 and August 2024 beyond those in the reviews.
Cite this review
Pith. "Pith review of Arabic Chatbot Technologies in Education: An Overview." pith.science (2026). https://pith.science/paper/UBKYFYPL
@misc{pith2026250904066,
author = {Pith},
title = {Pith review of: Arabic Chatbot Technologies in Education: An Overview},
year = {2026},
howpublished = {\url{https://pith.science/paper/UBKYFYPL}},
note = {Machine review of arXiv:2509.04066}
}
read the original abstract
The recent advancements in Artificial Intelligence (AI) in general, and in Natural Language Processing (NLP) in particular, and some of its applications such as chatbots, have led to their implementation in different domains like education, healthcare, tourism, and customer service. Since the COVID-19 pandemic, there has been an increasing interest in these digital technologies to allow and enhance remote access. In education, e-learning systems have been massively adopted worldwide. The emergence of Large Language Models (LLM) such as BERT (Bidirectional Encoder Representations from Transformers) and GPT (Generative Pre-trained Transformers) made chatbots even more popular. In this study, we present a survey on existing Arabic chatbots in education and their different characteristics such as the adopted approaches, language variety, and metrics used to measure their performance. We were able to identified some research gaps when we discovered that, despite the success of chatbots in other languages such as English, only a few educational Arabic chatbots used modern techniques. Finally, we discuss future directions of research in this field.
Reference graph
Works this paper leans on
-
[1]
https://doi.org/10.1016/j.caeai.2021.100033
Okonkwo, C.W., Ade-Ibijola, A.: Chatbots applications in education: A systematic review, (2021). https://doi.org/10.1016/j.caeai.2021.100033
-
[2]
https://doi.org/10.1145/3447735
Darwish, K., Habash, N., Abbas, M., Al-Khalifa, H., Al-Natsheh, H.T., Bouamor, H., Bouzoubaa, K., Cavalli-Sforza, V., El-Beltagy, S.R., El-Hajj, W., Jarrar, M., Mubarak, H.: A panoramic survey of natural language processing in the Arab world, (2021). https://doi.org/10.1145/3447735
-
[3]
https://doi.org/10.1016/j.jksuci.2019.02.006
Guellil, I., Saâdane, H., Azouaou, F., Gueni, B., Nouvel, D.: Arabic natural language processing: An overview, https://www.sciencedirect.com/science/article/pii/S1319157818310553, (2021). https://doi.org/10.1016/j.jksuci.2019.02.006
-
[4]
Alazzam, B.A., Alkhatib, M., Shaalan, K.: Arabic Educational Neural Network Chat-bot. Information Sciences Letters. 12, (2023). https://doi.org/10.18576/isl/120654
-
[5]
Computer Methods and Programs in Biomedicine Update
Ahmed, A., Ali, N., Alzubaidi, M., Zaghouani, W., Abdalrazaq, A., Househ, M.: Arabic chatbot technologies: A scoping review. Computer Methods and Programs in Biomedicine Update. 2, (2022). https://doi.org/10.1016/j.cmpbup.2022.100057
-
[6]
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A....
work page 2020
-
[7]
In: Advances in Intelligent Systems and Computing (2021)
El Hefny, W., Mansy, Y., Abdallah, M., Abdennadher, S.: Jooka: A Bilingual Chatbot for University Admission. In: Advances in Intelligent Systems and Computing (2021). https://doi.org/10.1007/978-3-030-72660-7_64
-
[8]
International Journal of BOURHIL , B., y EL YOUNOUSSI, Y.(2024)
Alqahtani, Q., Alrwais, O.: Building a Machine Learning Powered Chatbot for KSU Blackboard Users. International Journal of BOURHIL , B., y EL YOUNOUSSI, Y.(2024). Arabic Chatbot Technologies in Education: An Overview . En C. Rusu et al., (1ª ed.), T ransform ación digital en la educación: innovaciones y desafíos desde los cam pus virtuales (pp. 11-14). Hu...
-
[9]
International Journal of Advanced Computer Science and Applications
Alsheddi, A.S., Alhenaki, L.S.: English and Arabic Chatbots: A Systematic Literature Review. International Journal of Advanced Computer Science and Applications. 13, (2022). https://doi.org/10.14569/IJACSA.2022.0130876
-
[10]
In: Lecture Notes in Engineering and Computer Science (2013)
Alobaidi, O.G., Crockett, K.A., O’Shea, J.D., Jarad, T.M.: Abdullah: An intelligent arabic conversational tutoring system for modern islamic education. In: Lecture Notes in Engineering and Computer Science (2013)
work page 2013
-
[11]
Alobaidi, O.G., Smieee, K.A.C., Mieee, J.D.O.S., Jarad, T.M.: The application of learning theories into Abdullah: An intelligent Arabic conversational agent tutor. In: ICAART 2015 - 7th International Conference on Agents and Artificial Intelligence, Proceedings (2015). https://doi.org/10.5220/0005197003610369
-
[12]
Hijjawi, M., Bandar, Z., Crockett, K., McLean, D.: ArabChat: An arabic conversational agent. In: 2014 6th International Conference on Computer Science and Information Technology, CSIT 2014 - Proceedings (2014). https://doi.org/10.1109/CSIT.2014.6806005
-
[13]
International Journal of Advanced Computer Science and Applications
Hijjawi, M., Qattous, H., Alsheiksalem, O.: Mobile Arabchat: An Arabic Mobile-Based Conversational Agent. International Journal of Advanced Computer Science and Applications. 6, (2015). https://doi.org/10.14569/ijacsa.2015.061016
-
[14]
International Journal of Advanced Computer Science and Applications
Hijjawi, M., Bandar, Z., Crockett, K.: The Enhanced Arabchat: An Arabic Conversational Agent. International Journal of Advanced Computer Science and Applications. 7, (2016). https://doi.org/10.14569/ijacsa.2016.070247
-
[15]
Aljameel, S.S., O’Shea, J.D., Crockett, K.A., Latham, A., Kaleem, M.: Development of an Arabic Conversational Intelligent Tutoring System for Education of children with ASD. In: 2017 IEEE International Conference on Computational Intelligence and Virtual Environments for Measurement Systems and Applications, CIVEMSA 2017 - Proceedings (2017). https://doi....
-
[16]
In: Advances in Intelligent Systems and Computing (2019)
Aljameel, S., O’Shea, J., Crockett, K., Latham, A., Kaleem, M.: LANAI: An Arabic Conversational Intelligent Tutoring System for Children with ASD. In: Advances in Intelligent Systems and Computing (2019). https://doi.org/10.1007/978-3-030-22871-2_34
-
[17]
Journal of Applied Computer Science & Mathematics
ALMURTADHA, Y.: LABEEB: Intelligent Conversational Agent Approach to Enhance Course Teaching and Allied Learning Outcomes attainment. Journal of Applied Computer Science & Mathematics. 13, (2019). https://doi.org/10.4316/jacsm.201901001
-
[18]
Elgibreen, H., Almazyad, S., Shuail, S. Bin, Qahtani, M. Al, Alhwiseen, L.: Robot Framework for Anti-Bullying in Saudi Schools. In: Proceedings - 4th IEEE International Conference on Robotic Computing, IRC 2020 (2020). https://doi.org/10.1109/IRC.2020.00033
-
[19]
Computer Applications in Engineering Education
Sweidan, S.Z., Abu Laban, S.S., Alnaimat, N.A., Darabkh, K.A.: SIAAAC: A student interactive assistant android application with chatbot during COVID-19 pandemic. Computer Applications in Engineering Education. 29, (2021). https://doi.org/10.1002/cae.22419
-
[20]
In: 2021 9th International Conference on Information and Education Technology, ICIET 2021 (2021)
Sweidan, S.Z., Abu Laban, S.S., Alnaimat, N.A., Darabkh, K.A.: SEG-COVID: A Student Electronic Guide within Covid-19 Pandemic. In: 2021 9th International Conference on Information and Education Technology, ICIET 2021 (2021). https://doi.org/10.1109/ICIET51873.2021.9419656. BOURHIL , B., y EL YOUNOUSSI, Y.(2024). Arabic Chatbot Technologies in Education: A...
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.