Pith. sign in

REVIEW 3 major objections 6 minor 68 references

Educators' Perceptions of Large Language Models as Tutors: Comparing Human and AI Tutors in a Blind Text-only Setting

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Educators rate an LLM tutor above human tutors on all four teaching-quality metrics in a blind text-only comparison.

desk verdict A genuinely new blind comparison of LLM vs human tutoring on latent quality metrics, but the abstract overstates the result: engagement is not significant and the human baseline is a specific low-effort crowdworker corpus, not 'human tutors' in general. read the letter →

arxiv 2506.08702 v1 pith:27BOY4PJ submitted 2025-06-10 cs.ET

classification cs.ET
keywords LLMtutorhumancomparisonblindpreferenceannotationteachingqualitymetricsengagementempathyscaffoldingconcisenessMathDialMWPTutorwordproblems
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks teachers to compare snippets of human and LLM tutoring conversations without knowing which is which. On grade-school math word problems, the teachers rate the LLM tutor as better than the human tutor on engagement, empathy, scaffolding, and conciseness. The LLM's advantage is statistically significant for conciseness, empathy, and scaffolding; the engagement gap falls just short of significance (p=0.09). The authors argue this suggests LLMs can take over repetitive tutoring duties and reduce the load on human teachers, while cautioning that the human tutor baseline comes from paid crowdworkers who knew their student was simulated, which may not reflect real-world human tutoring.

What carries the argument

The central object is the blind pairwise preference comparison setup: annotators see a math word problem and two five-turn conversation snippets side by side, with left/right positions randomized, and choose 'Left is Better,' 'Right is Better,' or 'Both are Equal' on each of four metrics. This setup is intentionally designed to isolate the latent, textually observable qualities of tutoring—engagement, empathy, scaffolding, conciseness—from learning outcomes, which prior comparisons of LLM and human tutors have focused on.

What would settle it

A replication study with experienced, motivated human tutors (e.g., credentialed teachers paid per successful session and told the student is real) comparing their text-only tutoring against MWPTutor on the same 210 problems; if the human tutors match or beat the LLM on empathy, scaffolding, and conciseness, the paper's central claim that LLMs outperform human tutors in text-only tutoring would be falsified.

Watch

Extended reading notes

Core claim

In a blind, text-only preference task with 35 annotators who have teaching experience, annotators rate MWPTutor—an LLM-based tutor with guardrails—as better than human tutors from the MathDial dataset on all four metrics: engagement, empathy, scaffolding, and conciseness. Across 210 conversation pairs, the mean preference score (on a scale from -5, all humans, to +5, all LLM) is 0.55 for conciseness, 0.65 for empathy, 0.55 for scaffolding, and 0.25 for engagement, with effect sizes of 0.25, 0.23, 0.22, and 0.09 respectively. The empathy gap is the largest, with 80% of annotators preferring the LLM tutor more often than the human tutor. The paper also reports that LLMs' own judgments align with the human preference direction but are far stronger, and that correlations between human and LLM metric scores are weak, indicating imperfect alignment between what LLMs and humans perceive as good tutoring.

Load-bearing premise

The findings rest on the assumption that the MathDial human-tutor conversations are a fair representation of how human tutors actually teach, even though those tutors were crowdworkers paid a fixed small amount who knew their student was an AI and had no incentive to perform well.

Editorial extensions

If this is right

  • If replicated in more natural settings with real students and teachers, the results suggest LLMs can handle text-based tutoring roles for well-defined subjects like grade-school math, freeing human teachers for mentoring and socio-emotional duties.
  • The finding that humans prefer the LLM's conciseness and scaffolding despite its longer conversations implies that perceived progress, not utterance count, drives judgments of conciseness—a signal for how to design future tutor interfaces.
  • The gap between LLMs' strong self-preferences and humans' moderate preferences indicates that LLM-based evaluation of tutoring quality is not yet a reliable substitute for human judgment, pointing to a need for alignment work.
  • The study's methodological template—parallel human/LLM tutoring corpora with simulated students, blind comparison, and four subjectively defined metrics—can be extended to other subjects and tutor designs where comparable datasets exist.
  • The authors' observation that tutors who expressed scaffolding intent but implemented it poorly were rated worse suggests that intent annotations in tutoring datasets do not guarantee perceived pedagogical quality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The strongest direct corollary the authors do not fully spell out: if crowdworkers—who had no performance incentive and knew the student was simulated—still underperform an LLM on empathy and scaffolding, deployed human tutors under real workload and burnout pressures may show an even larger gap in text-only interactions, strengthening the case for LLM delegation but also raising equity concerns f
  • A testable extension: present the same conversation pairs to annotators with no teaching experience. The paper recruits only teaching-experienced annotators; if the preference ordering persists in naive judges, the effect is about general conversational quality rather than pedagogical expertise, which would change the interpretation.
  • A dangerous implication left implicit: the study compares a guardrailed LLM tutor (MWPTutor) against unguarded humans, but the Appendix shows a bare GPT-4o, which the authors found makes factual errors, still wins decisively on these subjective metrics. That suggests perceived quality and factual correctness can diverge, and future tutor deployments may need explicit correctness guardrails even wh
  • The truncation to five turns means the comparison captures only the opening moves of tutoring; extending the snippet window to ten or fifteen turns could test whether the LLM's advantage persists as the conversation deepens into sustained scaffolding.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper reports a blind, text-only comparison in which human annotators with teaching experience rated 5-turn snippets from two tutoring corpora: MathDial, a corpus of human (Prolific crowdworker) tutors interacting with a simulated student, and MWPTutor, an LLM-based tutor with guardrails. Annotators chose which tutor was better on four subjective metrics: engagement, empathy, scaffolding, and conciseness. The authors report that annotators perceived the LLM tutor as better on all four metrics, with the largest advantage in empathy (80% of annotators preferring the LLM), and they supplement this with LLM self-judgments and several post-hoc analyses. The paper also releases the 210 annotated conversation pairs.

Significance. If the central claim were fully supported, this would be a valuable contribution to the ongoing debate about whether LLM-based tutors can match or exceed human tutors on perceived pedagogical qualities. The blind pairwise design, the use of experienced-teacher annotators, and the public release of the annotation data are practical strengths. The study also provides a useful cautionary result that LLM self-judgments are poorly aligned with human judgments. However, the central claim is currently overstated relative to the evidence, and the representativeness of the human baseline is the main threat to external validity.

major comments (3)
  1. [Abstract; §4.2, Table 1] The abstract states that annotators perceive LLMs as showing 'higher performance than human tutors in all 4 metrics,' but Table 1 shows that the advantage for Engagement is not statistically significant at the paper's own threshold (p = 0.09 > 0.01). The supported statement is that the LLM has higher point estimates on all four metrics and statistically significant advantages on Conciseness, Empathy, and Scaffolding only. The abstract and Section 5.4 should be revised to reflect this distinction.
  2. [§5.2; Limitations] The human baseline (MathDial) was collected under conditions that are likely to depress tutor performance: Prolific workers paid a fixed small amount, knowing the student was simulated, with no performance incentives. The authors explicitly acknowledge this in Section 5.2 ('Knowing that the student is in fact an AI ... might have contributed to the teachers not doing their best'). Because the central claim is that LLMs are perceived as better than human tutors, the representativeness of this baseline is load-bearing. The paper should either provide evidence that MathDial tutors are representative of typical human tutoring (e.g., by comparing effort or quality metrics to other human-tutoring corpora) or substantially weaken the conclusion to 'better than the specific MathDial corpus' and discuss the external-validity threat more prominently.
  3. [§3.1; §6] The comparison is between a single LLM tutor (MWPTutor, built by the authors' group and explicitly selected because it enforces correctness) and a single human-tutoring corpus (MathDial). The paper consistently uses the generic terms 'LLM' and 'human tutor' in the abstract and conclusions. This overgeneralization is not supported by the design. The claims should be framed as a system-specific comparison, or the paper should include at least one additional LLM tutor and one additional human-tutoring corpus to justify the general phrasing.
minor comments (6)
  1. [§3.3] There is a typo: 'thereafter, it we had 150 slides' should be 'thereafter, we had 150 slides.'
  2. [Figures 1 and 2] The color descriptions ('brightest red indicating minimum possible score of −3' etc.) are difficult to follow; please consider labeling the color scale directly on the figures or using a more standard legend.
  3. [Table 1] The table reports one-sided p-values, but the text does not justify why one-sided tests are appropriate. If the hypothesis is directional, please state this explicitly; otherwise, report two-sided p-values.
  4. [§4.4] The system name is spelled inconsistently as both 'MWPTutor' and 'MWPtutor' (e.g., in the Engagement paragraph). Please use a single consistent spelling.
  5. [§5.1] The reference to 'Fig. 4 in the Mathdial paper' is not self-contained; since the figure is not reproduced, please describe the relevant result in text.
  6. [Limitations] The first limitation bullet ends with 'this setting only' without a terminal period, and the final paragraph of the Limitations section has a run-on style; please proofread.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the preference results are direct human annotations, not reduced from prior equations or fitted parameters.

full rationale

This paper is an empirical measurement study rather than a derivation chain, so the circularity patterns do not apply. The central claim, that annotators with teaching experience perceive the LLM tutor as better on the four metrics, is supported by blind pairwise human annotations (Table 1) collected independently of the authors' prior models. The two corpora compared, MathDial and MWPTutor, are prior works with overlapping authorship, but they serve as experimental materials, not as premises that mathematically force the outcome. The comparison was genuinely open: for Engagement in the 'Fresh' scenario, MathDial came out on top (d = 0.30, p = 0.001), showing the result was not predetermined by construction. The LLM-as-judge results are reported separately, and the paper explicitly acknowledges their limitation: 'LLMs are likely to be biased towards LLM-generated text.' The main weaknesses identified by the reader, such as the representativeness of Prolific crowdworker tutors and the acknowledged possibility that 'Knowing that the student is in fact an AI ... might have contributed to the teachers not doing their best,' are threats to ecological validity and generalizability, not circularity. No fitted input is renamed as a prediction, no uniqueness theorem is imported from the authors, and no self-citation carries the argument's weight. The finding is a legitimate empirical observation, even if its external validity is debatable.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper is an empirical comparison, so it introduces no fitted parameters or new entities. Its central claim rests on domain assumptions about representativeness of the datasets and the validity of subjective annotation, several of which the authors themselves flag in the Limitations section.

assumptions (5)
  • domain assumption MathDial conversations represent human tutor behavior
    The entire comparison treats MathDial as the human side; the paper's Limitations question the ecological validity.
  • domain assumption Five-turn snippet suffices to judge the four metrics
    Section 3.3 truncates all dialogs to 5 turns and acknowledges this 'increases the epistemic noise of the task.'
  • domain assumption Prolific self-reported teaching experience is a valid proxy for educator status
    Section 3.4 notes Prolific does not verify qualifications, so some annotators may not be educators.
  • domain assumption Annotator pairwise preferences are a valid measure of tutoring quality
    The whole analysis rests on subjective judgments; the paper discusses low inter-annotator agreement.
  • domain assumption Simulated students (InstructGPT) produce realistic student behaviors
    Both datasets use the same simulated student, and the paper avoids real student data for comparability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Educators' Perceptions of Large Language Models as Tutors: Comparing Human and AI Tutors in a Blind Text-only Setting." pith.science (2026). https://pith.science/paper/27BOY4PJ

@misc{pith2026250608702,
  author       = {Pith},
  title        = {Pith review of: Educators' Perceptions of Large Language Models as Tutors: Comparing Human and AI Tutors in a Blind Text-only Setting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/27BOY4PJ}},
  note         = {Machine review of arXiv:2506.08702}
}
read the original abstract

The rapid development of Large Language Models (LLMs) opens up the possibility of using them as personal tutors. This has led to the development of several intelligent tutoring systems and learning assistants that use LLMs as back-ends with various degrees of engineering. In this study, we seek to compare human tutors with LLM tutors in terms of engagement, empathy, scaffolding, and conciseness. We ask human tutors to annotate and compare the performance of an LLM tutor with that of a human tutor in teaching grade-school math word problems on these qualities. We find that annotators with teaching experience perceive LLMs as showing higher performance than human tutors in all 4 metrics. The biggest advantage is in empathy, where 80% of our annotators prefer the LLM tutor more often than the human tutors. Our study paints a positive picture of LLMs as tutors and indicates that these models can be used to reduce the load on human teachers in the future.

Figures

Figures reproduced from arXiv: 2506.08702 by the authors.

Figure 1
Figure 1. Fractions of conversation pairs which received particular scores for each metric from LLMs. Scores increase left to right, with the brightest red indicating minimum possible score of −3, the dullest red indicating −1, grey indicating 0, the dullest green indicating +1 and the brightest green indicating the maximum possible score of +3 4 Results and Analysis We mentioned earlier that our metrics involve some scope fo… view at source ↗
Figure 2
Figure 2. Fractions of conversation pairs which received particular scores for each metric from humans. Scores increase left to right, with the brightest red indicating minimum possible score of −5, the dullest red indicating −1, grey indicating 0, the dullest green indicating +1 and the brightest green indicating the maximum possible score of +5 Conciseness(Human) Engagement(Human) Empathy(Human) Scaffolding(Human) Concisene… view at source ↗
Figure 3
Figure 3. Correlation between various metrics, as anno [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Fractions of conversation pairs which received particular scores for each metric from s. Scores increase [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Instructions for Evaluation Metrics 18 [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: Combined View of Intro Slide and Metric Evaluation Slide [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

68 extracted references · 38 canonical work pages

  1. [1]

    New survey finds students are replacing human tutors with chatgpt

    2023. New survey finds students are replacing human tutors with chatgpt. https://www.intelligent.com/new-survey-finds-students-are-replacing-human-tutors-with-chatgpt/

  2. [2]

    Digital education council global ai student survey

    2024. Digital education council global ai student survey. https://www.digitaleducationcouncil.com/post/digital-education-council-global-ai-student-survey-2024

  3. [3]

    The multifaceted roles of a teacher: Beyond the classroom

    2024. The multifaceted roles of a teacher: Beyond the classroom. https://teachers.institute/learning-teaching/roles-of-teacher-beyond-classroom/

  4. [4]

    Fabian Albers, Melanie Trypke, Ferdinand Stebner, Joachim Wirth, and Jan L. Plass. 2023. https://doi.org/10.1111/bjep.12592 Different types of redundancy and their effect on learning and cognitive load . British Journal of Educational Psychology, 93(S2):339--352

  5. [5]

    Haady Abdilnibi Altememy, Nour Raheem Neamah, Rabaa Mazhair, Nada Sami Naser, Ali Afrawi Fahad, Nazar Abdulghffar al sammarraie, Hidab Rasul Sharif, Mohamed Amer Alseidi, and Mohammed Yousif Oudah Al-Muttar. 2023. Ai tools' impact on student performance: Focusing on student motivation & engagement in iraq. Przestrze \'n Spo eczna (Social Space) , 23(2):143--165

  6. [6]

    JR Anderson, AT Corbett, KR Koedinger, and R Pelletier. 1995. https://doi.org/10.1207/s15327809jls0402 _2 Cognitive tutors: Lessons learned . Journal of the Learning Sciences, 4(2):167--207

  7. [7]

    Axelson and Arend Flick

    Rick D. Axelson and Arend Flick. 2010. https://doi.org/10.1080/00091383.2011.533096 Defining student engagement . Change: The Magazine of Higher Learning, 43(1):38--43

  8. [8]

    Fatemeh Bambaeeroo and Nasrin Shokrpour. 2017. The impact of the teachers’ non-verbal communication on success in teaching. Journal of advances in medical education & professionalism, 5(2):51

Show all 68 references
  1. [9]

    Benjamin S. Bloom. 1984. https://doi.org/10.3102/0013189X013006004 The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring . Educational Researcher, 13(6):4--16

  2. [10]

    Timothy B Bostic. 2014. Teacher empathy and its relationship to the standardized test scores of diverse secondary english students. Journal of Research in Education, 24(1):3--16

  3. [11]

    Laurie Butgereit, Herman Martinus, and Muna Mahmoud Abugosseisa. 2023. https://doi.org/10.1109/INES59282.2023.10297824 Prof pi: Tutoring mathematics in arabic language using gpt-4 and whatsapp . In 2023 IEEE 27th International Conference on Intelligent Engineering Systems (INE...

  4. [12]

    Andrew Caines, Helen Yannakoudakis, Helena Edmondson, Helen Allen, Pascual P \'e rez-Paredes, Bill Byrne, and Paula Buttery. 2020. https://aclanthology.org/2020.nlp4call-1.2/ The teacher-student chatroom corpus . In Proceedings of the 9th Workshop on NLP for Computer Assisted ...

  5. [13]

    C Daryl Cameron, Cendri A Hutcherson, Amanda M Ferguson, Julian A Scheffer, Eliana Hadjiandreou, and Michael Inzlicht. 2019. Empathy is hard work: People choose to avoid empathy because of its cognitive costs. Journal of Experimental Psychology: General, 148(6):962

  6. [14]

    Ching-Huei Chen and Ching-Ling Chang. 2024. Effectiveness of ai-assisted game-based learning on science learning outcomes, intrinsic motivation, cognitive load, and learning behavior. Education and Information Technologies, pages 1--22

  7. [15]

    Vuthea Chheang, Shayla Sharmin, Rommy Marquez-Hernandez, Megha Patel, Danush Rajasekaran, Gavin Caulfield, Behdokht Kiafar, Jicheng Li, Pinar Kullu, and Roghayeh Leila Barmaki. 2024. https://arxiv.org/abs/2306.17278 Towards anatomy education with generative ai-based virtual as...

  8. [16]

    Nandita Chitrakar and Dr P.M. 2023. https://doi.org/10.59828/ijsrmst.v2i11.158 Frustration and its influences on student motivation and academic performance . International Journal of Scientific Research in Modern Science and Technology, 2:01--09

  9. [17]

    Rudrajit Choudhuri, Dylan Liu, Igor Steinmacher, Marco Gerosa, and Anita Sarma. 2023. https://arxiv.org/abs/2312.11719 How far are we? the triumphs and trials of generative ai in learning software engineering . Preprint, arXiv:2312.11719

  10. [18]

    Sankalan Pal Chowdhury, Vil \'e m Zouhar, and Mrinmaya Sachan. 2024. https://api.semanticscholar.org/CorpusID:267657803 Autotutor meets large language models: A language model tutor with rich pedagogy and guardrails . In ACM Conference on Learning @ Scale

  11. [19]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://arxiv.org/abs/2110.14168 Training verifiers to solve math word problems . ...

  12. [20]

    Sidney D'Mello and Art Graesser. 2013. Autotutor and affective autotutor: Learning by talking with cognitively and emotionally intelligent computers that talk back. ACM Transactions on Interactive Intelligent Systems (TiiS), 2(4):1--39

  13. [21]

    Gerald A. Goldin. 2000. https://doi.org/10.1207/S15327833MTL0203_3 Affective pathways and representation in mathematical problem solving . Mathematical Thinking and Learning, 2(3):209--219

  14. [22]

    Sven Jacobs and Steffen Jaschke. 2024. https://doi.org/10.1109/educon60312.2024.10578838 Evaluating the application of large language models to generate feedback in programming education . In 2024 IEEE Global Engineering Education Conference (EDUCON). IEEE

  15. [23]

    Donna Ault Jacobson. 2016. Causes and effects of teacher burnout. Walden University

  16. [24]

    Slava Kalyuga and John Sweller. 2014. The Redundancy Principle in Multimedia Learning, page 247–262. Cambridge Handbooks in Psychology. Cambridge University Press

  17. [25]

    Argyro Kavadella, Marco Antonio Dias Da Silva, Eleftherios G Kaklamanos, Vasileios Stamatopoulos, Kostis Giannakopoulos, and 1 others. 2024. Evaluation of chatgpt’s real-life implementation in undergraduate dental education: mixed methods study. JMIR Medical Education, 10(1):e51344

  18. [26]

    Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley, Paul Denny, Michelle Craig, and Tovi Grossman. 2024. https://doi.org/10.1145/3613904.3642773 Codeaid: Evaluating a classroom deployment of an llm-based programming assistant that balances student and educat...

  19. [27]

    Kulik and James A

    Chen-Lin C. Kulik and James A. Kulik. 1991. https://doi.org/10.1016/0747-5632(91)90030-5 Effectiveness of computer-based instruction: An updated analysis . Computers in Human Behavior, 7(1):75--94

  20. [28]

    Benjamin Kutsyuruba and Lorraine Godden. 2019. The role of mentoring and coaching as a means of supporting the well-being of educators and students. International Journal of Mentoring and Coaching in Education, 8(4):229--234

  21. [29]

    Hao Lei, Yunhuo Cui, and Wenye Zhou. 2018. https://doi.org/10.2224/sbp.7054 Relationships between student engagement and academic achievement: A meta-analysis . Social Behavior and Personality: an international journal, 46:517--528

  22. [30]

    Wengxi Li, Roy Pea, Nick Haber, and Hari Subramonyam. 2024. https://arxiv.org/abs/2405.12946 Tutorly: Turning programming videos into apprenticeship learning environments with llms . Preprint, arXiv:2405.12946

  23. [31]

    Anna Lieb and Toshali Goel. 2024. Student interaction with newtbot: An llm-as-tutor chatbot for secondary physics education. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1--8

  24. [32]

    Mark Liffiton, Brad E Sheese, Jaromir Savelka, and Paul Denny. 2023. Codehelp: Using large language models with guardrails for scalable support in programming classes. In Proceedings of the 23rd Koli Calling International Conference on Computing Education Research, pages 1--11

  25. [33]

    Rongxin Liu, Carter Zenke, Charlie Liu, Andrew Holmes, Patrick Thornton, and David J. Malan. 2024. https://doi.org/10.1145/3626252.3630938 Teaching cs50 with ai: Leveraging generative artificial intelligence in computer science education . In Proceedings of the 55th ACM Techni...

  26. [34]

    Wenhan Lyu, Yimeng Wang, Tingting (Rachel) Chung, Yifan Sun, and Yixuan Zhang. 2024. https://doi.org/10.1145/3657604.3662036 Evaluating the effectiveness of llms in introductory computer science education: A semester-long field study . In Proceedings of the Eleventh ACM Confer...

  27. [35]

    Ross B MacDonald. 2000. The master tutor: A guidebook for more effective tutoring. Cambridge Stratford Study Skills Institute

  28. [36]

    Jakub Macina, Nico Daheim, Sankalan Chowdhury, Tanmay Sinha, Manu Kapur, Iryna Gurevych, and Mrinmaya Sachan. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.372 M ath D ial: A dialogue tutoring dataset with rich pedagogical properties grounded in math reasoning problems...

  29. [37]

    Tsediso Michael Makoelle. 2019. Teacher empathy: a prerequisite for an inclusive classroom. Encyclopedia of teacher education, 11(2):27--39

  30. [38]

    Kaushal Kumar Maurya, KV Srivatsa, Kseniia Petukhova, and Ekaterina Kochmar. 2024. Unifying ai tutor evaluation: An evaluation taxonomy for pedagogical ability assessment of llm-powered ai tutors. arXiv preprint arXiv:2412.09416

  31. [39]

    George A Miller. 1956. The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological review, 63(2):81

  32. [40]

    OpenAI and GPT4 Team. 2024. https://arxiv.org/abs/2303.08774 Gpt-4 technical report . Preprint, arXiv:2303.08774

  33. [41]

    Maciej Pankiewicz and Ryan S. Baker. 2024. https://doi.org/10.1145/3649217.3653608 Navigating compiler errors with ai assistance - a study of gpt hints in an introductory programming course . In Proceedings of the 2024 on Innovation and Technology in Computer Science Education...

  34. [42]

    Pardos and Shreya Bhandari

    Zachary A. Pardos and Shreya Bhandari. 2024. https://doi.org/10.1371/journal.pone.0304013 Chatgpt-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills . PLOS ONE, 19(5):1--18

  35. [43]

    Minju Park, Sojung Kim, Seunghyun Lee, Soonwoo Kwon, and Kyuseok Kim. 2024. https://doi.org/10.1145/3613905.3651122 Empowering personalized learning through a conversation-based tutoring system with student modeling . In Extended Abstracts of the CHI Conference on Human Factor...

  36. [44]

    Graesser

    Natalie Person and Arthur C. Graesser. 2002. Human or computer? autotutor in a bystander turing test. In Intelligent Tutoring Systems, pages 821--830, Berlin, Heidelberg. Springer Berlin Heidelberg

  37. [45]

    Abey P Philip and Dawn Bennett. 2021. Using deliberate mistakes to heighten student attention. Journal of University Teaching and Learning Practice, 18(6):193--212

  38. [46]

    Petra Polakova and Blanka Klimova. 2024. https://doi.org/10.1080/2331186X.2024.2355385 Implementation of ai-driven technology into education – a pilot study on the use of chatbots in foreign language learning . Cogent Education, 11(1):2355385

  39. [47]

    Laryn Qi, J. D. Zamfirescu-Pereira, Taehan Kim, Björn Hartmann, John DeNero, and Narges Norouzi. 2024. https://arxiv.org/abs/2406.05603 A knowledge-component-based methodology for evaluating ai assistants . Preprint, arXiv:2406.05603

  40. [48]

    Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025. https://arxiv.org/abs/2412.15115 Qwe...

  41. [49]

    Robin Schmucker, Meng Xia, Amos Azaria, and Tom Mitchell. 2023. Ruffle&riley: Towards the automated induction of conversational tutoring systems. arXiv preprint arXiv:2310.01420

  42. [50]

    Sleeman and J.S

    D. Sleeman and J.S. Brown. 1982. https://books.google.ch/books?id=pjqcAAAAMAAJ Intelligent Tutoring Systems . Computers and people series. Academic Press

  43. [51]

    Adam Smith. 2006. Cognitive empathy and emotional empathy in human behavior and evolution. The Psychological Record, 56(1):3--21

  44. [52]

    Katherine Stasaski, Kimberly Kao, and Marti A. Hearst. 2020. https://doi.org/10.18653/v1/2020.bea-1.5 CIMA : A large open access dialogue dataset for tutoring . In Proceedings of the Fifteenth Workshop on Innovative Use of NLP for Building Educational Applications, pages 52--6...

  45. [53]

    Snežana Stojiljković, Gordana Djigić, and Blagica Zlatković. 2012. https://doi.org/10.1016/j.sbspro.2012.12.021 Empathy and teachers’ roles . Procedia - Social and Behavioral Sciences, 69:960--966. International Conference on Education & Educational Psychology (ICEEPSY 2012)

  46. [54]

    Mayeesha Tabassum and Md Mahbub-ul Alam. 2024. Exploring multifaceted expectations from teachers: An analysis from guardians’ and students’ perspective. Indonesian Journal of Social Research (IJSR), 6(2):141--155

  47. [55]

    Maung Thway, Jose Recatala-Gomez, Fun Siong Lim, Kedar Hippalgaonkar, and Leonard W. T. Ng. 2024. https://arxiv.org/abs/2406.07796 Battling botpoop using genai for higher education: A study of a retrieval augmented generation chatbots impact on learning . Preprint, arXiv:2406.07796

  48. [56]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://arxiv.org/abs/2302.13971 Llama:...

  49. [57]

    A. M. Turing. 1950. https://doi.org/10.1093/mind/LIX.236.433 I.—computing machinery and intelligence . Mind, LIX(236):433--460

  50. [58]

    Kurt VanLehn. 2011. https://doi.org/10.1080/00461520.2011.611369 The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems . Educational Psychologist, 46(4):197--221

  51. [59]

    Alessandro Vanzo, Sankalan Pal Chowdhury, and Mrinmaya Sachan. 2024. Gpt-4 as a homework tutor can improve student engagement and learning outcomes. arXiv preprint arXiv:2409.15981

  52. [60]

    Robert J Walker. 2008. Twelve characteristics of an effective teacher: A longitudinal, qualitative, quasi-research study of in-service and pre-service teachers' opinions. educational HORIZONS, pages 61--68

  53. [61]

    Chiu, Jiayin Zhi, Shaun M

    Ruiyi Wang, Stephanie Milani, Jamie C. Chiu, Jiayin Zhi, Shaun M. Eack, Travis Labrum, Samuel M. Murphy, Nev Jones, Kate Hardy, Hong Shen, Fei Fang, and Zhiyu Zoey Chen. 2024. https://arxiv.org/abs/2405.19660 Patient- : Using large language models to simulate patients for trai...

  54. [62]

    David Wood, Jerome S Bruner, and Gail Ross. 1976. The role of tutoring in problem solving. Journal of child psychology and psychiatry, 17(2):89--100

  55. [63]

    Bissyandé, and Shunfu Jin

    Boyang Yang, Haoye Tian, Weiguo Pian, Haoran Yu, Haitao Wang, Jacques Klein, Tegawendé F. Bissyandé, and Shunfu Jin. 2024. https://arxiv.org/abs/2406.13972 Cref: An llm-based conversational software repair framework for programming tutors . Preprint, arXiv:2406.13972

  56. [64]

    Xiajun Yu, Changkang Sun, Binghai Sun, Xuhui Yuan, Fujun Ding, and Mengxie Zhang. 2022. https://doi.org/10.3390/su14106071 The cost of caring: Compassion fatigue is a special form of teacher burnout . Sustainability, 14(10)

  57. [65]

    Lingling Zang and Yameng Chen. 2022. Relationship between person-organization fit and teacher burnout in kindergarten: the mediating role of job satisfaction. Frontiers in Psychiatry, 13:948934

  58. [66]

    Jiayue Zhang, Yiheng Liu, Wenqi Cai, Lanlan Wu, Yali Peng, Jingjing Yu, Senqing Qi, Taotao Long, and Bao Ge. 2024. https://arxiv.org/abs/2403.16687 Investigation of the effectiveness of applying chatgpt in dialogic teaching using electroencephalography . Preprint, arXiv:2403.16687

  59. [67]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  60. [68]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.