REVIEW 3 major objections 5 minor 71 references
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that zero-shot grade prompting does not make GPT-4 explanations match their intended educational level, with a 50% perceived-match rate versus 79% for lay human-curated explanations, and that users find the model…
desk verdict A useful benchmark and two careful human studies, but the headline gap rests on an author-curated baseline that inflates the comparison; the core finding that zero-shot grade prompting is insufficient still holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the ELI-WHY benchmark plus two human-study protocols. ELI-WHY's 13,392 "Why" questions provide a common stimulus across disciplines; the perceived-background-match task asks raters to judge which educational level an explanation suits; and the informativeness task asks learners whether an explanation introduces a new concept and connects it to prior knowledge. Alongside these, the paper uses readability and surface-form metrics (Flesch-Kincaid Reading Ease, Linsear Write, Dale-Chall, sentence counts, reading time, and a complex-word ratio) and introduces the term interpretation collapse for the observation that score distributions for elementary, high-school, and graduate explanations map onto overlapping interpreted grade ranges. The manually web-retrieved lay explanations serve as the human baseline against which both the 50%-versus-79% match gap and the 20% informativeness gap are defined.
What would settle it
A concrete check: re-run both human studies with lay-curated explanations produced by independent crowdworkers who do not know the intended grade level, using the same 40 questions. If the independent baseline's match rate drops to near GPT-4's 50%, or if users rate it no more informative than GPT-4, the central claim that lay humans outperform GPT-4 in grade tailoring would be falsified.
Extended reading notes
Core claim
The paper's central discovery is that prompting a model for a grade level does not reliably produce an explanation that educators or learners experience as being at that grade level. In the educator study, only about 50% of GPT-4 explanations generated for elementary, high-school, or graduate users were perceived as matching their intended level, versus 79.16% for manually web-retrieved lay explanations. In the learner study, GPT-4 explanations were rated roughly 20% less informative on average, because they were less likely both to introduce new concepts and to connect those concepts to prior knowledge; the shortfall was largest among graduate and high-school users. Automated readability metrics tell the same story: explanations grow longer and use more complex words as the target level rises, yet the interpreted U.S. grade levels of the three audiences overlap heavily, so the outputs are not meaningfully distinguished. The paper concludes that prompting alone is insufficient for audience adaptation and that pedagogical utility must be assessed by human judgments of fit and informativeness.
Load-bearing premise
The load-bearing premise is that the manually web-retrieved explanations curated by the paper's authors are a fair and unbiased representation of lay human explanations; if that baseline is tilted toward text that obviously matches grade labels, both headline gaps are inflated.
Editorial extensions
If this is right
- If correct, zero-shot grade prompting is not enough to tailor explanations; model providers should not assume a prompt like "explain to an elementary-school student" produces the intended register.
- The ELI-WHY benchmark and its two human protocols give educators and developers a way to measure pedagogical fit before deploying explanations.
- The gap is largest for advanced learners, making graduate-level users the group least well served by current explanations.
- Automated readability metrics can serve only as a partial screening signal, and then only when interpreted through grade-level mappings rather than raw scores.
- A relative 20% informativeness deficit means users relying on model explanations are systematically getting less new, connectable information than they would from curated web explanations.
Reading between the lines
- One implication the authors leave implicit is that the same two protocols could benchmark multi-turn or retrieval-augmented explanation systems; only zero-shot single-turn generation is tested here, so the measured failure may be a floor, not a ceiling.
- A testable extension: run the educator and learner studies on models fine-tuned or preference-optimized for audience fit; if their match rate and informativeness remain near GPT-4's levels, the gap reflects a deeper representational limit rather than a prompting shortfall.
- The paper's case study about questions resembling popular simplified explanations suggests a more general inference: questions with a strong association to one audience in training data may systematically pull generation toward that audience, which could be probed computationally by regressing perceived match on similarity to audience-labeled corpora.
- The use of three coarse educational levels as a proxy for informational needs is a simplification; a finer-grained measure of prior knowledge could plausibly change the magnitude of the reported gaps.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ELI-Why, a benchmark of 13,392 'Why' questions, and uses it to evaluate whether language models can tailor explanations to three educational backgrounds (elementary, high school, graduate). Explanations are generated from four model families using zero-shot grade prompts; the main human studies use GPT-4 and compare against two baselines: a web-retrieved baseline (Google Featured Snippets) and an author-curated 'Manually Web-Retrieved' set of lay human explanations. Study 1 asks raters, in the role of educators, to identify the perceived educational background of each explanation; Study 2 asks participants, in the role of learners, whether each explanation introduces new concepts that connect to their prior knowledge. The paper reports a perceived background match of roughly 50% for GPT-4 versus 79% for the manual baseline, and a 20% informativeness gap, alongside an 'interpretation collapse' in readability metrics across models. The authors conclude that zero-shot grade prompting is insufficient for pedagogical tailoring and that ELI-Why provides a way to measure this failure.
Significance. This paper addresses a timely and practically important problem: whether prompt-based grade tailoring is sufficient for LLM explanations in education. Its design has several strengths: the benchmark is large and diverse across STEM and non-STEM domains; the human studies use role-based tasks (educator versus learner) with natural-language justifications; the automated analysis covers four model families and multiple readability and surface-form metrics; the ELI5-similarity case study offers a concrete, testable hypothesis for the observed mismatches; and the authors state that they will release the benchmark and all generated explanations. If the headline gaps survive a fairer baseline construction and proper uncertainty quantification, the paper would provide a useful evaluation resource and a clear negative result about zero-shot prompting for grade-adaptive explanation. However, the headline magnitudes rest on the author-curated baseline and on small subsets without reported confidence intervals, so the size of the effect is not yet fully supported.
major comments (3)
- [§3 (Baseline explanations), §4.1, §5] The headline comparisons in §4.1 (50% vs 79% perceived background match) and §5 (20% informativeness gap) both use the 'Manually Web-Retrieved' baseline described in §3. This baseline is curated by the authors with full knowledge of the intended grade level: the protocol explicitly searches ELI5 for elementary explanations, blogs for high school, and journals/papers for graduate explanations. Such a label-aware selection procedure is an oracle rather than a representative sample of lay human explanations; it is likely to favor texts with conspicuous grade-level markers, which could inflate both the 79% match rate and the 20% informativeness advantage. The limitations section addresses prompt strategy and subset size but not the construction of this baseline. I request either a blinded or independently curated baseline (curators who do not know the study's hypothesis, even if they must know the target grade), a pre-existing corpus with natural grade metadata, or at minimum a sensitivity analysis showing the results under a relaxed curation protocol.
- [§5 (Human study and metrics, Results)] The '20% more informative' result is reported without statistical inference or a clear definition of the aggregate. The study uses 40 questions and 200 responses per background; Figure 6 shows no confidence intervals or error bars. It is not specified whether the 20% is an average over the three backgrounds, over the two metrics (% Informative and % Matched Informative), or some other aggregation; the text switches between '20% less suited' (abstract) and 'relatively 20% more informative' (§5). Moreover, the design selects only GPT-4 explanations that were perceived to match their intended background (§5: 'we use GPT-4-generated explanations that were perceived to match their intended educational backgrounds'), while no such selection is described for the Manually Web-Retrieved baseline; if the baseline was not equally filtered, the comparison is not apples-to-apples. Please provide the aggregation formula, a confidence interval or significance test, and a clear statement of whether the same inclusion criterion was applied to both explanation sets.
- [§4.1 (User study design and Results)] The perceived background match study reports aggregate match rates (about 50% for GPT-4 and 79.16% for Manually Web-Retrieved) without inter-annotator agreement statistics or confidence intervals. The control baseline is based on only 40 questions and three annotations per question; with n=40, a 95% confidence interval for a 79.16% proportion is roughly 63%–90%, which overlaps with plausible values for the GPT-4 condition. Without agreement measures (e.g., Fleiss' kappa) or interval estimates, it is difficult to assess whether the gap is as large as claimed. Please report these quantities or justify why a majority vote of three annotations is sufficient to support the headline comparison.
minor comments (5)
- [Abstract] The word 'effectivenes' appears to be a typo for 'effectiveness', and the benchmark name is inconsistent throughout ('ELI-Why' in the title and abstract, 'ELI-WHY' elsewhere).
- [§4.1] The phrase 'GPT-4-generated rationales' introduces an unnecessary synonym; elsewhere in the paper the terms 'explanations' is used consistently, and using both terms can confuse readers.
- [Figure 6] The bar chart in Figure 6 has no error bars or confidence intervals; adding them would help the reader assess the reliability of the reported differences, especially given the small number of questions.
- [§B.2] In the filtering description, 'Annotator’s were posed the following question' should be 'Annotators were posed the following question'.
- [§E.4] The caption for Figure 15 contains the typo 'dispalus' for 'displays'; please correct it.
Circularity Check
No significant circularity: headline gaps rest on external human ratings and standard readability formulas, not on a fitted or self-referential derivation.
full rationale
This paper contains no derivation chain that reduces to its own inputs. The central measurements—the 50% vs. 79% perceived background match and the 20% informativeness gap—are produced by external human raters acting as educators and learners, and by standard readability formulas (Flesch-Kincaid, Linsear Write, Dale-Chall), not by any fitted parameter or by the model being tested. The ELI-Why questions are GPT-4-generated, so the evaluation target is partly self-referential, but GPT-4 does not score or select its own explanations; the judgments come from human participants and independent text metrics. The 'Manually Web-Retrieved' baseline is curated by the authors with knowledge of the intended grade (Section 3), which could inflate the reported gap; this is a validation and confound concern that would require blinded curation or a pre-existing corpus with natural grade metadata, but it is not a circular step because no equation or fitted value is reused as a prediction. Self-citations such as Joshi et al. (2023) and Zhao et al. (2024) appear in related work and are not load-bearing for the headline claims. The limitations section acknowledges the GPT-4-generated benchmark and the subset size, but not the baseline construction; that omission affects interpretability of the gap magnitude rather than establishing circularity.
Assumptions & free parameters
assumptions (7)
- domain assumption Highest educational degree attained is a valid proxy for a learner's informational needs.
- domain assumption The three discrete levels Elementary, High School, and Graduate are "least overlapping" and sufficient to cover users.
- domain assumption Manually Web-Retrieved explanations curated by the authors are representative of lay human explanations.
- domain assumption Zero-shot prompting with role instructions is a fair way to elicit grade-tailored explanations.
- domain assumption ELI-WHY questions, overgenerated by GPT-4 and filtered by answerability, are a valid sample of real pedagogical "Why" questions.
- domain assumption Crowdworkers' perceived educational background and informativeness judgments are valid measures of pedagogical utility.
- standard math Standard statistical tests (Mann-Whitney U, KS, correlation ratio) apply appropriately to explanation score distributions.
Cite this review
Pith. "Pith review of ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations." pith.science (2026). https://pith.science/paper/T4DGOO6Y
@misc{pith2026250614200,
author = {Pith},
title = {Pith review of: ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations},
year = {2026},
howpublished = {\url{https://pith.science/paper/T4DGOO6Y}},
note = {Machine review of arXiv:2506.14200}
}
read the original abstract
Language models today are widely used in education, yet their ability to tailor responses for learners with varied informational needs and knowledge backgrounds remains under-explored. To this end, we introduce ELI-Why, a benchmark of 13.4K "Why" questions to evaluate the pedagogical capabilities of language models. We then conduct two extensive human studies to assess the utility of language model-generated explanatory answers (explanations) on our benchmark, tailored to three distinct educational grades: elementary, high-school and graduate school. In our first study, human raters assume the role of an "educator" to assess model explanations' fit to different educational grades. We find that GPT-4-generated explanations match their intended educational background only 50% of the time, compared to 79% for lay human-curated explanations. In our second study, human raters assume the role of a learner to assess if an explanation fits their own informational needs. Across all educational backgrounds, users deemed GPT-4-generated explanations 20% less suited on average to their informational needs, when compared to explanations curated by lay people. Additionally, automated evaluation metrics reveal that explanations generated across different language model families for different informational needs remain indistinguishable in their grade-level, limiting their pedagogical effectiveness.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Maxime Adolphe, Marion Pech, Masataka Sawayama, Denis Maurel, Alexandra Delmas, Pierre-Yves Oudeyer, and Hélène Sauzéon. 2023. https://doi.org/10.31234/osf.io/2wg59 Exploring the potential of artificial intelligence in individualized cognitive training: a systematic review
-
[2]
Tal August, Kyle Lo, Noah A. Smith, and Katharina Reinecke. 2024. https://doi.org/10.1145/3613904.3642289 Know your audience: The benefits and pitfalls of generating plain language summaries beyond the "general" audience . In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI '24, New York, NY, USA. Association for Computing...
arXiv 2024
-
[3]
Hearst, Andrew Head, and Kyle Lo
Tal August, Lucy Lu Wang, Jonathan Bragg, Marti A. Hearst, Andrew Head, and Kyle Lo. 2023. https://doi.org/10.1145/3589955 Paper plain: Making medical research papers approachable to healthcare consumers with natural language processing . ACM Trans. Comput.-Hum. Interact., 30(5)
doi:10.1145/3589955 2023
-
[4]
Astrid Bertrand, Tiphaine Viard, Rafik Belloum, James R Eagan, and Winston Maxwell. 2023. On selective, mutable and dialogic xai: A review of what users say about different types of interactive explanations. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1--21
work page 2023
-
[5]
Maureen A Callanan and Lisa M Oakes. 1992. Preschoolers' questions and parents' explanations: Causal thinking in everyday activity. Cognitive development, 7(2):213--233
work page 1992
-
[6]
Serina Chang, Ashton Anderson, and Jake M. Hofman. 2025. https://arxiv.org/abs/2504.07114 Chatbench: From static benchmarks to human-ai evaluation . Preprint, arXiv:2504.07114
arXiv 2025
-
[7]
Hanjie Chen, Faeze Brahman, Xiang Ren, Yangfeng Ji, Yejin Choi, and Swabha Swayamdipta. 2023. https://doi.org/10.18653/v1/2023.acl-long.112 REV : Information-theoretic evaluation of free-text rationales . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2007--2030, Toronto, Canada. A...
-
[8]
Inyoung Cheong, King Xia, K. J. Kevin Feng, Quan Ze Chen, and Amy X. Zhang. 2024. https://arxiv.org/abs/2402.01864 (a)i am not a lawyer, but...: Engaging legal experts towards responsible llm policies for legal advice . Preprint, arXiv:2402.01864
work page Pith review arXiv 2024
Show all 71 references
-
[9]
Alexis Chevalier, Jiayi Geng, Alexander Wettig, Howard Chen, Sebastian Mizera, Toni Annala, Max Jameson Aragon, Arturo Rodr \' guez Fanlo, Simon Frieder, Simon Machado, et al. 2024. Language models as science tutors. arXiv preprint arXiv:2402.11111
2024 arXiv
-
[10]
Michelle M Chouinard, Paul L Harris, and Michael P Maratsos. 2007. Children's questions: A mechanism for cognitive development. Monographs of the society for research in child development, pages i--129
2007
-
[11]
why does rain fall?
Kathleen H Corriveau and Katelyn E Kurkul. 2014. “why does rain fall?”: Children prefer to learn from an informant who uses noncircular explanations. Child development, 85(5):1827--1835
2014
-
[12]
Edgar Dale and Jeanne S Chall. 1948. A formula for predicting readability: Instructions. Educational research bulletin, pages 37--54
1948
-
[13]
Huw C Davies, Rebecca Eynon, and Cory Salveson. 2021. https://doi.org/10.1177/0038038520967888 The mobilisation of ai in education: A bourdieusean field analysis . Sociology, 55(3):539--560
2021 doi
-
[14]
Vera Demberg and Frank Keller. 2008. Data from eye-tracking corpora as evidence for theories of syntactic processing complexity. Cognition, 109(2):193--210
2008
-
[15]
John H Falk and Mark D Needham. 2013. Factors contributing to adult knowledge of science and technology. Journal of Research in Science Teaching, 50(4):431--452
2013
-
[16]
Nils Feldhus, Aliki Anagnostopoulou, Qianli Wang, Milad Alshomary, Henning Wachsmuth, Daniel Sonntag, and Sebastian M \"o ller. 2024. Towards modeling and evaluating instructional explanations in teacher-student dialogues. In Proceedings of the 2024 International Conference on...
2024
-
[17]
Ronald Aylmer Fisher. 1970. Statistical methods for research workers. In Breakthroughs in statistics: Methodology and distribution, pages 66--70. Springer
1970
-
[18]
R Flesch. 1948. A new readability yardstick. J. Appl. Psychol., 32(3):221--233
1948
-
[19]
Rudolf Flesch. 1979. How to write plain english. University of Canterbury. Available at http://www. mang. canterbury. ac. nz/writing\_guide/writing/flesch. shtml.[Retrieved 5 February 2016]
1979
-
[20]
Lukas H \"o per and Carsten Schulte. 2024. New perspectives on the future of computing education: Teaching and learning explanatory models. In Proceedings of the 24th Koli Calling International Conference on Computing Education Research, pages 1--8
2024
-
[21]
Yi-Sheng Hsu, Nils Feldhus, and Sherzod Hakimov. 2024. Free-text rationale generation under readability level control. arXiv preprint arXiv:2407.01384
2024 arXiv
-
[22]
Brihi Joshi, Ziyi Liu, Sahana Ramnath, Aaron Chan, Zhewei Tong, Shaoliang Nie, Qifan Wang, Yejin Choi, and Xiang Ren. 2023. https://doi.org/10.18653/v1/2023.acl-long.392 Are machine rationales (not) useful to humans? measuring and improving human utility of free-text rationale...
2023 doi
-
[23]
Irina Jurenka, Markus Kunesch, Kevin R. McKee, Daniel Gillick, Shaojian Zhu, Sara Wiltberger, Shubham Milind Phal, Katherine Hermann, Daniel Kasenberg, Avishkar Bhoopchand, Ankit Anand, Miruna Pîslar, Stephanie Chan, Lisa Wang, Jennifer She, Parsa Mahmoudieh, Aliya Rysbek, Wei...
2024
-
[24]
Mert Karabacak and Konstantinos Margetis. 2023. Embracing large language models for medical applications: opportunities and challenges. Cureus, 15(5)
2023
-
[25]
u chemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan G \
Enkelejda Kasneci, Kathrin Se ler, Stefan K \"u chemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan G \"u nnemann, Eyke H \"u llermeier, et al. 2023. Chatgpt for good? on opportunities and challenges of large language models for education....
2023
-
[26]
Deborah Kelemen. 1999. Why are rocks pointy? children's preference for teleological explanations of th natural world. Developmental psychology, 35(6):1440
1999
-
[27]
Minsun Kim, SeonGyeom Kim, Suyoun Lee, Yoosang Yoon, Junho Myung, Haneul Yoo, Hyunseung Lim, Jieun Han, Yoonsu Kim, So-Yeon Ahn, et al. 2024 a . Designing prompt analytics dashboards to analyze student-chatgpt interactions in efl writing. arXiv preprint arXiv:2405.19691
2024 arXiv
-
[28]
Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2024 b . Evallm: Interactive evaluation of large language model prompts on user-defined criteria. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1--21
2024
-
[29]
Fishburne, Richard L
Peter Kincaid, Robert P. Fishburne, Richard L. Rogers, and Brad S. Chissom. 1975. https://api.semanticscholar.org/CorpusID:61131325 Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel
1975
-
[30]
David A Kolb et al. 2007. The Kolb learning style inventory. Hay Resources Direct Boston, MA
2007
-
[31]
Katelyn E Kurkul and Kathleen H Corriveau. 2018. Question, explanation, follow-up: A mechanism for learning from others? Child Development, 89(1):280--294
2018
-
[32]
Seongyun Lee, Sue Hyun Park, Seungone Kim, and Minjoon Seo. 2024. https://arxiv.org/abs/2405.17977 Aligning to thousands of preferences via system message generalization . Preprint, arXiv:2405.17977
2024 arXiv
-
[33]
Yoonjoo Lee, Tae Soo Kim, Sungdong Kim, Yohan Yun, and Juho Kim. 2023. Dapie: Interactive step-by-step explanatory dialogues to answer children’s why and how questions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1--22
2023
-
[34]
is chatgpt a better explainer than my professor?
Grace Li, Milad Alshomary, and Smaranda Muresan. 2024. " is chatgpt a better explainer than my professor?": Evaluating the explanation capabilities of llms in conversation compared to a human baseline. arXiv preprint arXiv:2406.18512
2024 arXiv
-
[35]
Smith, and Yejin Choi
Alisa Liu, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.508 WANLI : Worker and AI collaboration for natural language inference dataset creation . In Findings of the Association for Computational Linguistics: EMNLP 202...
2022 doi
-
[36]
Tania Lombrozo and Nicholas Z Gwynne. 2014. Explanation and inference: Mechanistic and functional explanations guide property generalization. Frontiers in Human Neuroscience, 8:700
2014
-
[37]
Silvia B Lovato, Anne Marie Piper, and Ellen A Wartella. 2019. Hey google, do unicorns exist? conversational agents as a path to answers to children's questions. In Proceedings of the 18th ACM international conference on interaction design and children, pages 301--313
2019
-
[38]
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2...
2022
-
[39]
Guoqing Luo, Yu Tong Han, Lili Mou, and Mauajama Firdaus. 2023. Prompt-based editing for text style transfer. arXiv preprint arXiv:2301.11997
2023 arXiv
-
[40]
Henry B Mann and Donald R Whitney. 1947. On a test of whether one of two random variables is stochastically larger than the other. The annals of mathematical statistics, pages 50--60
1947
-
[41]
Amanda M McCarthy and Frank C Keil. 2023. A right way to explain? function, mechanism, and the order of explanations. Cognition, 238:105494
2023
-
[42]
Chancharik Mitra, Mihran Miroyan, Rishi Jain, Vedant Kumud, Gireeja Ranade, and Narges Norouzi. 2024. Retllm-e: Retrieval-prompt strategy for question-answering on student discussion forums. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 232...
2024
-
[43]
Randall Munroe. 2015. Thing explainer: complicated stuff in simple words. Hachette UK
2015
-
[44]
Sanghee Oh, Jung Sun Oh, and Chirag Shah. 2008. The use of information sources by internet users in answering questions. Proceedings of the American Society for Information Science and Technology, 45(1):1--13
2008
-
[45]
John O'hayre. 1966. Gobbledygook has gotta go. US Department of the Interior, Bureau of Land Management
1966
-
[46]
Ben Prystawski, Michael Li, and Noah Goodman. 2023. Why think step by step? reasoning emerges from the locality of experience. Advances in Neural Information Processing Systems, 36:70926--70947
2023
-
[47]
Romain Puech, Jakub Macina, Julia Chatain, Mrinmaya Sachan, and Manu Kapur. 2024. https://arxiv.org/abs/2410.03781 Towards the pedagogical steering of large language models for tutoring: A case study with modeling productive failure . Preprint, arXiv:2410.03781
2024 arXiv
-
[48]
Rod D Roscoe and Michelene TH Chi. 2008. Tutor learning: The role of explaining and responding to questions. Instructional science, 36:321--350
2008
-
[49]
Alexis Ross and Jacob Andreas. 2024. https://doi.org/10.18653/v1/2024.acl-long.718 Toward in-context teaching: Adapting examples to students' misconceptions . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pa...
2024 doi
-
[50]
Adena Schachner, Liqi Zhu, Jing Li, and Deborah Kelemen. 2017. Is the bias for function-based explanations culturally universal? children from china endorse teleological explanations of natural phenomena. J. Exp. Child Psychol., 157:29--48
2017
-
[51]
Robin Schmucker, Meng Xia, Amos Azaria, and Tom Mitchell. 2024. Ruffle&riley: Insights from designing and evaluating a large language model-based conversational tutoring system. In International Conference on Artificial Intelligence in Education, pages 75--90. Springer
2024
-
[52]
Woosuk Seo, Chanmo Yang, and Young-Ho Kim. 2024. Chacha: Leveraging large language models to prompt children to share their emotions about personal events. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1--20
2024
-
[53]
Maja Stahl, Leon Biermann, Andreas Nehring, and Henning Wachsmuth. 2024. Exploring llm prompting strategies for joint essay scoring and feedback generation. arXiv preprint arXiv:2404.15845
2024 arXiv
-
[54]
John Stamper, Ruiwei Xiao, and Xinying Hou. 2024. Enhancing llm-based feedback: Insights from intelligent tutoring systems and the learning sciences. In International Conference on Artificial Intelligence in Education, pages 32--43. Springer
2024
-
[55]
Wadim Strielkowski, Veronika Grebennikova, Alexander Lisovskiy, Guzalbegim Rakhimova, and Tatiana Vasileva. 2024. Ai-driven adaptive learning for sustainable educational transformation. Sustainable Development
2024
-
[56]
Justin Sulik, Jeroen van Paridon, and Gary Lupyan. 2023. Explanations in the wild. Cognition, 237(105464):105464
2023
-
[57]
Lihui Sun and Liang Zhou. 2024. https://doi.org/10.1177/07356331241277937 Does generative artificial intelligence improve the academic achievement of college students? a meta-analysis . Journal of Educational Computing Research, 62(7):1896--1933
2024 doi
-
[58]
Siddharth Suri, Scott Counts, Leijie Wang, Chacha Chen, Mengting Wan, Tara Safavi, Jennifer Neville, Chirag Shah, Ryen W White, Reid Andersen, et al. 2024. The use of generative search engines for knowledge work and complex tasks. arXiv preprint arXiv:2404.04268
2024 arXiv
-
[59]
Clemmie Telford. 2021. But Why?: How to answer tricky questions from kids and have an honest conversation with yourself. Headline Home
2021
-
[60]
Barbara Tizard and Martin Hughes. 2008. Young children learning. John Wiley & Sons
2008
-
[61]
Ahmed Tlili, Boulus Shehata, Michael Agyemang Adarkwah, Aras Bozkurt, Daniel T Hickey, Ronghuai Huang, and Brighter Agyemang. 2023. What if the devil is my guardian angel: Chatgpt as a case study of using chatbots in education. Smart learning environments, 10(1):15
2023
-
[62]
Bailin Wang, Zi Wang, Xuezhi Wang, Yuan Cao, Rif A Saurous, and Yoon Kim. 2023. Grammar prompting for domain-specific language generation with large language models. Advances in Neural Information Processing Systems, 36:65030--65055
2023
-
[63]
Adrian F Ward. 2021. People mistake the internet’s knowledge for their own. Proceedings of the National Academy of Sciences, 118(43):e2105061118
2021
-
[64]
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652
2021 arXiv
-
[65]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837
2022
-
[66]
Lyumanshan Ye, Jiandong Jiang, Danni Chang, and Pengfei Liu. 2024. Storypark: Leveraging large language models to enhance children story learning through child-ai collaboration storytelling. arXiv preprint arXiv:2405.06495
2024
-
[67]
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2023. Evaluating large language models at evaluating instruction following. arXiv preprint arXiv:2310.07641
2023 arXiv
-
[68]
Chao Zhang, Xuechen Liu, Katherine Ziska, Soobin Jeon, Chi-Lin Yu, and Ying Xu. 2024. Mathemyths: leveraging large language models to teach mathematical language through child-ai co-creative storytelling. In Proceedings of the CHI Conference on Human Factors in Computing Syste...
2024
-
[69]
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024. https://arxiv.org/abs/2405.01470 Wildchat: 1m chatgpt interaction logs in the wild . Preprint, arXiv:2405.01470
2024 arXiv
-
[70]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[71]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.