REVIEW 4 major objections 4 minor 72 references
Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education
T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper claims that making trustworthiness visible—as five metrics and traceable visualizations—raises expert agreement when learning engineers evaluate LLM responses, and argues for hybrid human-in-the-loop evaluation.
desk verdict A well-executed co-design study with a plausible but under-supported central claim: the violation pipeline is unvalidated, so the observed IRR increase could be anchoring on shared noise. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a per-response violation pipeline built on claim extraction and textual entailment. Each response's claims are checked against the textbook passage, the learner's summary, the chat history, and the latest learner message; entailment labels of entails/contradicts/neutral determine whether a claim counts as misinformation, hallucination, sycophancy, and so on—with neutral treated as the signature of an unsupported claim. Learning-method measures are scored by an LLM-as-a-judge against explicit pedagogical principles, and response-quality measures come from validated conversational features. The visualizations then map each flagged span back into the response text,
What would settle it
Run the same tournament with the same interfaces but with metric flags computed on shuffled or randomly mislabeled responses. If expert agreement rises as much with meaningless flags, the effect comes from having any shared display, not from trustworthy signal. Also, compare the automated flags to expert human judgments on the same 30 responses: near-chance agreement would directly undermine the claim.
Extended reading notes
Core claim
The central claim is that operationalizing LLM trustworthiness as a set of violation metrics plus text-traceable visualizations improves the reliability of expert evaluation of LLM responses in education. The authors co-constructed five metrics—Truthfulness, Diplomacy, Disposition, Learning Methods, and Response Quality—composed of 20 measures, each computed per response by an automated pipeline. They then ran an LLM prompt tournament in which 12 learning engineers chose between paired responses, half the time with the metrics and visualizations visible. Expert agreement was higher with metrics visible (alpha 0.4931 vs 0.3987), and the distribution of 'best' responses shifted toward lower-vi
Load-bearing premise
The whole finding rests on the assumption that the automated flags—built from claim extraction, textual entailment, an LLM judge, and fitted thresholds—are accurate enough to be trusted as evidence; the paper states their accuracy in educational contexts is not yet tested.
Editorial extensions
If this is right
- If the claim holds, prompt-evaluation panels in education can raise agreement without imposing a single weighting of criteria.
- Automated violation flags can serve as a first-pass screening layer, with engineers resolving ambiguous cases.
- The metric set offers a reusable starting point for education-specific LLM evaluation that can be adjusted to other pedagogical contexts.
- Comparison-oriented visualizations (summary overview, glyphs, donuts) are the features engineers actually use; single-metric encodings like opacity and underlining need refinement.
- Evaluation tools that show trustworthiness alongside rubric criteria can change which prompt wins, since the rubric winner may be the least trustworthy.
Reading between the lines
- The agreement gain may come less from new information than from a shared reference point: any consistent, inspectable flag set might produce part of the effect, and a condition with plain-text metric summaries would isolate what visualization adds.
- The same co-designed metric set could plausibly be adapted to other high-stakes LLM domains where experts disagree on outputs, since the mechanism is general rather than education-specific.
- Because the study used one model and one dialogue task, the durability of the effect under repeated use and across models is open; a longitudinal replication would clarify whether trust in the measures grows or decays.
- The dataset-derived thresholds (e.g., Participation, Diversity) should be recalibrated per corpus; without that, transferred deployments risk spurious flags that could erode the agreement benefit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a three-phase co-design study with learning engineers building an LLM-powered digital textbook. The authors first develop five trustworthiness metrics (Truthfulness, Diplomacy, Disposition, Learning Methods, Response Quality) comprising 20 measures, then design visualizations that map metric violations onto LLM responses, and finally evaluate the tools in an A/B prompt-evaluation tournament with 12 learning engineers rating 30 match-ups. The headline result is that making trustworthiness metrics visible increased inter-rater reliability (Krippendorff's alpha from 0.3987 without metrics to 0.4931 with metrics) and helped participants reconcile conflicting rubric objectives. The paper also contributes design guidelines for future LLM evaluation tools.
Significance. If the headline effect holds, this is a useful contribution to HCI and learning-engineering practice: it provides a concrete, domain-grounded operationalization of LLM trustworthiness for educational evaluation and a careful design-study template. The study has real strengths: it is ecologically valid (real student conversations, authentic prompt templates, within-subjects alternating interface conditions), it is longitudinal and co-design-based, and the authors are explicit about several key limitations—notably that the automated measures' accuracy in educational contexts is untested (§8.3) and that Participation and Diversity thresholds were fitted to the study dataset (§5.3). The qualitative findings—metrics acting as a procedural checklist, visualizations supporting trade-off reasoning—are credible and well supported by participant quotes. However, the central quantitative claim currently rests on an unvalidated measurement pipeline and an unsupported statistical comparison, which are the main barriers to accepting the paper's conclusions.
major comments (4)
- [§5.3, §8.3] The central claim presupposes that the 20 measures produce valid violation flags that learning engineers can use as trustworthy evidence. Section 8.3 states that 'their accuracy applied to educational contexts remains untested,' and no precision/recall, expert-label agreement, or any other validation is reported for any measure. The observed IRR increase (α0=0.3987 to α1=0.4931) is therefore also consistent with raters anchoring on a shared but possibly arbitrary overlay rather than on genuine trustworthiness information. Please add a validation study—e.g., expert labels on a held-out sample of responses with per-measure agreement, or at minimum a sensitivity analysis showing that the IRR effect survives when flagged-but-questionable responses are removed.
- [§7.1] The primary quantitative claim is the difference between α0=0.3987 and α1=0.4931. No confidence interval, bootstrap, or significance test is reported for either alpha or their difference, despite §4.3.6 adopting bootstrapped 95% CIs and the 'CIs do not overlap' rule. Both values are also below the conventional 0.67 reliability threshold, so the increase is from 'poor/moderate' to 'still poor/moderate.' Please report bootstrap CIs for each condition and a proper test of the difference (e.g., bootstrap percentile or permutation test); if the CIs overlap, the headline claim should be weakened accordingly.
- [§5.3 (Fig. 10Q,R)] The Participation and Diversity violation thresholds were, in the authors' own words, 'set from the distribution of our own dataset rather than an external standard' and 'require recalibration before being applied to another corpus.' Because the same dataset is used in the tournament, the 'With Metrics' condition displays flags that are partly fitted to the evaluation set. This is a circularity: the increase in agreement is not yet evidence for trustworthiness metrics in general, only for this dataset-calibrated instantiation. Please report thresholds derived from prior work or via cross-validation, and test the robustness of the IRR difference to reasonable threshold variations.
- [§4.3.3] Four of the twelve tournament raters participated in the earlier co-design phases for the metrics and visualizations (and, per §4.3.3, the additional eight were recruited later). Although participants had not seen the final visualizations, the co-designers helped define the design goals and measures; their prior investment may create demand characteristics or heightened fluency with the tool that the newly recruited raters lack. Please report whether the IRR increase (α0 vs. α1) holds when restricting the analysis to the eight non-co-design participants, or explicitly discuss this confound in the limitations.
minor comments (4)
- [§7.2 / Fig. 16] The time-on-task comparison is reported descriptively (median 79.55s vs. 62.95s; 6/12 slower). This is fine as an exploratory result, but it would fit the paper's own §4.3.6 statistical framework to report bootstrap CIs for the medians or means.
- [Fig. 17 / Fig. 19] The 'Ratio of (higher−lower)-violation responses picked' values are shown with CIs in Fig. 17, but Fig. 19 is presented without CIs or raw counts. Adding counts would help readers assess the strength of the 'selected in spite of violation' patterns.
- [Fig. 10O] The Dale-Chall readability formula and its constants are given only inside the figure. Since this is one of the 20 measures, please also define it in the text or a numbered equation so readers can reproduce the computation.
- [General] The CCS Concepts line still contains 'Do Not Use This Code' placeholders; this should be corrected before publication.
Circularity Check
No significant circularity; the central IRR finding is an empirical comparison, not a derivation, though dataset-fitted thresholds and an author-overlapping citation warrant noting.
full rationale
The paper's central claim—that making trustworthiness visible increased inter-rater reliability—is supported by a within-subjects prompt tournament experiment, not by a derivation from the paper's own equations. The disclosed fitting of Participation and Diversity thresholds 'from the distribution of our own dataset' (Section 5.3) is a validity limitation: the flags shown to raters are partly dataset-specific, but the outcome measure (IRR) is an empirical behavioral result, not a quantity forced by those thresholds. The paper explicitly states that 'their accuracy applied to educational contexts remains untested' (Section 8.3), which is an honest caveat about the pipeline rather than a circular reduction. The prompt tournament method is cited to the authors' prior work [27], but the method is fully described in Sections 3.1.2 and 4.3, so the citation is not load-bearing. The co-designers later serving as evaluators creates a reflexivity concern, but the IRR metric is computed from blind A/B choices and the paper acknowledges such individual-difference factors. Under the hard rule requiring an explicit reduction (e.g., Eq. X = Eq. Y by construction, or a fitted parameter renamed as a prediction), no such circular step can be quoted from this paper.
Assumptions & free parameters
free parameters (7)
- Entailment confidence threshold =
50%
- Semantic similarity threshold =
0.5 (text) / 0.7 (figure example)
- Certainty violation threshold =
< 5.5
- Readability violation threshold =
Dale-Chall score >= 0.6
- Subjectivity violation threshold =
> 0.5
- Participation violation threshold =
Gini > 0.4
- Diversity violation threshold =
topic vector score > 0.7
assumptions (5)
- domain assumption The automated violation pipeline produces flags accurate enough to support expert decisions
- domain assumption Textual entailment and semantic similarity models transfer to learner-generated text and textbook passages
- domain assumption LLM-as-a-Judge returns valid violations for Socratic, Scaffolding, Connections, Metacognitive, and SERT
- domain assumption The 12 learning engineer raters are representative of the broader learning engineer population
- domain assumption The alternating within-subjects interface design controls for learning and ordering effects
Cite this review
Pith. "Pith review of Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education." pith.science (2026). https://pith.science/paper/NZ6UC77Q
@misc{pith2026260804006,
author = {Pith},
title = {Pith review of: Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education},
year = {2026},
howpublished = {\url{https://pith.science/paper/NZ6UC77Q}},
note = {Machine review of arXiv:2608.04006}
}
read the original abstract
LLMs are reshaping educational technology, yet evaluating their responses for pedagogical alignment remains underexplored, relying heavily on the expertise of learning engineers building the technology. To bridge this gap, we explore trustworthiness as a structured lens for evaluation, leveraging existing measures of LLM trustworthiness to systematically identify potential pedagogical disruptions. Through a longitudinal co-design process with learning engineers developing an LLM-powered digital textbook, we: (1) co-constructed five trustworthiness metrics comprising 20 measures tailored to pedagogical use; (2) designed visualizations that map trustworthiness violations onto LLM responses; and (3) evaluated how these tools help learning engineers make A/B comparisons of LLM responses. Making trustworthiness explicit increased inter-rater reliability while helping learning engineers resolve conflicting objectives and produce more consistent judgments. We discuss the emergent benefits of trustworthiness as a lens for evaluating LLMs in education and propose new design guidelines for future evaluation tools that enable pedagogically-aligned, LLM-powered learning tools.
Figures
Figures from the paper (19 more)
Reference graph
Works this paper leans on
-
[1]
Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent Anti-Muslim Bias in Large Language Models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society(Virtual Event, USA)(AIES ’21). Association for Computing Machinery, New York, NY, USA, 298–306. doi:10.1145/3461702.3462624
-
[2]
Saleh Afroogh, Ali Akbari, Emmie Malone, Mohammadali Kargar, and Hananeh Alambeigi. 2024. Trust in AI: progress, challenges, and future directions.Humanities and Social Sciences Communications11, 1 (2024), 1568
work page 2024
- [3]
-
[4]
Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, et al. 2020. Toward trustworthy AI development: mechanisms for supporting verifiable claims.arXiv preprint arXiv:2004.07213(2020)
arXiv 2020
-
[5]
Davide Ceneda, Theresia Gschwandtner, Thorsten May, Silvia Miksch, Hans-Jörg Schulz, Marc Streit, and Christian Tominski. 2017. Characterizing Guidance in Visual Analytics.IEEE Transactions on Visualization and Computer Graphics 23, 1 (2017), 111–120. doi:10.1109/TVCG.2016.2598468
arXiv 2017
-
[6]
Jeanne S. Chall and Edgar Dale. 1995.Readability Revisited: The New Dale-Chall Readability Formula. Brookline Books, Cambridge, MA, USA
work page 1995
- [7]
-
[8]
A. Chatzimparmpas, R. M. Martins, I. Jusufi, K. Kucher, F. Rossi, and A. Kerren. 2020. The State of the Art in Enhancing Trust in Machine Learning Models with the Use of Visualizations.Computer Graphics Forum39, 3 (2020), 713–756. doi:10.1111/cgf.14034
Show all 72 references
-
[9]
Lijia Chen, Pingping Chen, and Zhijian Lin. 2020. Artificial Intelligence in Education: A Review.IEEE Access8 (2020), 75264–75278. doi:10.1109/ACCESS.2020.2988510
2020
-
[10]
Adam Coscia and Alex Endert. 2024. KnowledgeVIS: Interpreting Language Models by Comparing Fill-in-the-Blank Prompts.IEEE Transactions on Visualization and Computer Graphics30, 9 (2024), 6520–6532. doi:10.1109/TVCG.2023. 3346713
2024 doi
-
[11]
Choi, Scott Crossley, and Alex Endert
Adam Coscia, Langdon Holmes, Wesley Morris, Joon S. Choi, Scott Crossley, and Alex Endert. 2024. iScore: Visual Analytics for Interpreting How Language Models Automatically Score Summaries. InProceedings of the 2024 IUI Conference on Intelligent User Interfaces(Greenville, SC,...
2024
-
[12]
Adam J Coscia, Shunan Guo, Eunyee Koh, and Alex Endert. 2025. OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST ’25). Associati...
2025
-
[13]
Cotton Debby R
Peter A. Cotton Debby R. E. Cotton and J. Reuben Shipway. 2024. Chatting and cheating: Ensuring academic integrity in the era of ChatGPT.Innovations in Education and Teaching International61, 2 (2024), 228–239. arXiv:https://doi.org/10.1080/14703297.2023.2190148 doi:10.1080/14...
2024
-
[14]
2018.Learning engineering for online education
Chris Dede, Bror Saxberg, and John Richards. 2018.Learning engineering for online education. Vol. 95. Routledge New York
2018
-
[15]
Andre del Carpio Gutierrez, Paul Denny, and Andrew Luxton-Reilly. 2024. Automating Personalized Parsons Problems with Customized Contexts and Concepts(ITiCSE 2024). Association for Computing Machinery, New York, NY, USA, 688–694. doi:10.1145/3649217.3653568
2024
-
[16]
2016.Fair Statistical Communication in HCI
Pierre Dragicevic. 2016.Fair Statistical Communication in HCI. Springer International Publishing, Cham, 291–330. doi:10.1007/978-3-319-26633-6_13
2016 doi
-
[17]
Jessica Díaz, Jorge Pérez, Carolina Gallardo, and Ángel González-Prieto. 2023. Applying Inter-Rater Reliability and Agreement in collaborative Grounded Theory studies in software engineering.Journal of Systems and Software195 (2023), 111520. doi:10.1016/j.jss.2022.111520
2023
-
[18]
Evans and William Revelle
Anthony M. Evans and William Revelle. 2008. Survey and behavioral measurements of interpersonal trust.Journal of Research in Personality42, 6 (2008), 1585–1593. doi:10.1016/j.jrp.2008.07.011
2008 doi
-
[19]
Niles, Ken Pathak, and Steven Sloan
Md Meftahul Ferdaus, Mahdi Abdelguerfi, Elias Loup, Kendall N. Niles, Ken Pathak, and Steven Sloan. 2026. Towards Trustworthy AI: A Review of Ethical and Robust Large Language Models.ACM Comput. Surv.58, 7, Article 176 (Jan. 2026), 43 pages. doi:10.1145/3777382
2026 doi
-
[20]
Yingqiang Ge, Shuchang Liu, Zuohui Fu, Juntao Tan, Zelong Li, Shuyuan Xu, Yunqi Li, Yikun Xian, and Yongfeng Zhang. 2024. A Survey on Trustworthy Recommender Systems.ACM Trans. Recomm. Syst.3, 2, Article 13 (Nov. 2024), 68 pages. doi:10.1145/3652891
2024 doi
-
[21]
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2020, Trevor Cohn, Yulan He, and Yang Liu (Eds.)...
2020 doi
-
[22]
Kummerfeld, and Elena L
Katy Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K. Kummerfeld, and Elena L. Glassman. 2024. Supporting Sensemaking of Large Language Model Outputs at Scale. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Associa...
2024
-
[23]
Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A Survey on Automated Fact-Checking.Transactions of the Association for Computational Linguistics10 (2022), 178–206. doi:10.1162/tacl_a_00454
2022 doi
-
[24]
Silva, and Huamin Qu
Jianben He, Xingbo Wang, Shiyi Liu, Guande Wu, Cláudio T. Silva, and Huamin Qu. 2025. POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models. In2025 IEEE 18th Pacific Visualization Conference (PacificVis). 36–46. doi:10.1109/PACIFICVI...
2025
-
[25]
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. InProceedings of the 9th International Conference on Learning Representations (ICLR ’21). arXiv:2006.03654
2021 arXiv
-
[26]
Fred Hohman, Minsuk Kahng, Robert Pienta, and Duen Horng Chau. 2019. Visual Analytics in Deep Learning: An Interrogative Survey for the Next Frontiers.IEEE Transactions on Visualization and Computer Graphics25, 8 (2019), 2674–2693. doi:10.1109/TVCG.2018.2843369
2019
-
[27]
Langdon Holmes, Adam Coscia, Scott Crossley, Joon Suh Choi, and Wesley Morris. 2026. LLM Prompt Evaluation for Educational Applications.arXiv preprint arXiv:2601.16134(2026)
2026
-
[28]
Santos, Mercedes T
Wayne Holmes, Kaska Porayska-Pomsta, Ken Holstein, Emma Sutherland, Toby Baker, Simon Buckingham Shum, Olga C. Santos, Mercedes T. Rodrigo, Mutlu Cukurova, Ig Ibert Bittencourt, and Kenneth R. Koedinger. 2022. Ethics of AI in Education: Towards a Community-Wide Framework.Inter...
2022 doi
-
[29]
Xinlan Emily Hu. 2025. team_comm_tools: A Python Toolkit for Exploring the Communication Space. PsyArXiv preprint. doi:10.31234/osf.io/pz98q_v2
2025 doi
-
[30]
Yue Huang et al. 2024. Position: TrustLLM: Trustworthiness in Large Language Models. InForty-first International Conference on Machine Learning. https://openreview.net/forum?id=bWUU0LwwMp
2024
-
[31]
Bill Hunter. 2015. Teaching for engagement: Part 1–Constructivist principles, case-based teaching, and active learning. College Quarterly18, 2 (2015), n2. J. ACM, Vol. 37, No. 4, Article 111. Publication date: August 2026. 111:46 Coscia et al
2015
-
[32]
Muhammad Imran and Norah Almusharraf. 2024. Google Gemini as a next generation AI educational tool: a review of emerging educational technology.Smart Learning Environments11, 1 (23 May 2024), 22. doi:10.1186/s40561-024-00310-z
2024 doi
-
[33]
Ellen Jiang, Kristen Olson, Edwin Toh, Alejandra Molina, Aaron Donsbach, Michael Terry, and Carrie J. Cai. 2022. PromptMaker: Prompt-based Prototyping with Large Language Models. InExtended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (CHI EA ’22)...
2022
-
[34]
Minsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu, James Wexler, Emily Reif, Krystal Kallarackal, Minsuk Chang, Michael Terry, and Lucas Dixon. 2024. LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models. InExtended Abstracts of th...
2024
-
[35]
Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al. 2023. ChatGPT for good? On opportunities and challenges of large language models for education.Learning and...
2023
-
[36]
Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2024. EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Associat...
2024
-
[37]
Malcolm S. Knowles. 1978. Andragogy: Adult Learning Theory in Perspective.Community College Review5, 3 (1978), 9–20. doi:10.1177/009155217800500302
1978 doi
-
[38]
Katerina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé, Danai Myrtzani, Theodoros Evgeniou, Ion Androut- sopoulos, and John Pavlopoulos. 2025. Evaluation and Facilitation of Online Discussions in the LLM Era: A Survey. In Proceedings of the 2025 Conference on Empirical ...
2025 doi
-
[39]
Moritz Laurer, Wouter van Atteveldt, Andreu Casas, and Kasper Welbers. 2024. Less Annotating, More Classifying: Addressing the Data Scarcity Issue of Supervised Machine Learning with Deep Transfer Learning and BERT-NLI. Political Analysis32, 1 (2024), 84–100. doi:10.1017/pan.2023.20
2024 doi
-
[40]
Maxime Lelièvre, Amy Waldock, Meng Liu, Natalia Valdés Aspillaga, Alasdair Mackintosh, María José Ogando Portela, Jared Lee, Paul Atherton, Robin AA Ince, and Oliver GB Garrod. 2025. Benchmarking the pedagogical knowledge of large language models.arXiv preprint arXiv:2506.18710(2025)
2025 arXiv
-
[41]
Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou. 2023. Trustworthy AI: From Principles to Practices.ACM Comput. Surv.55, 9, Article 177 (Jan. 2023), 46 pages. doi:10.1145/3555803
2023 doi
-
[42]
Jakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur, Iryna Gurevych, and Mrinmaya Sachan. 2025. MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors.arXiv preprint arXiv:2502.18940 (2025)
2025
-
[43]
McNamara
Danielle S. McNamara. 2004. SERT: Self-Explanation Reading Training.Discourse Processes38, 1 (2004), 1–30. doi:10.1207/s15326950dp3801_1
2004 doi
-
[44]
Meyer, Ryan J
Jesse G. Meyer, Ryan J. Urbanowicz, Patrick C. N. Martin, Karen O’Connor, Ruowang Li, Pei-Chen Peng, Tiffani J. Bright, Nicholas Tatonetti, Kyoung Jae Won, Graciela Gonzalez-Hernandez, and Jason H. Moore. 2023. ChatGPT and large language models in academia: opportunities and c...
2023 doi
-
[45]
Mihran Miroyan, Chancharik Mitra, Rishi Jain, Gireeja Ranade, and Narges Norouzi. 2025. Analyzing Pedagogical Quality and Efficiency of LLM Responses with TA Feedback to Live Student Questions. InProceedings of the 56th ACM Technical Symposium on Computer Science Education V. ...
2025
-
[47]
Wesley Morris, Scott Crossley, Langdon Holmes, Chaohua Ou, Mihai Dascalu, and Danielle McNamara. 2025. Formative Feedback on Student-Authored Summaries in Intelligent Textbooks Using Large Language Models.International Journal of Artificial Intelligence in Education35, 3 (01 S...
2025 doi
-
[48]
Wesley Morris, Scott Crossley, Langdon Holmes, Chaohua Ou, Danielle McNamara, and Mihai Dascalu. 2023. Using Large Language Models to Provide Formative Feedback in Intelligent Textbooks. InArtificial Intelligence in Education. Posters and Late Breaking Results, Workshops and T...
2023
-
[49]
Wesley Morris, Langdon Holmes, Joon Suh Choi, and Scott Crossley. 2025. Automated Scoring of Constructed Response Items in Math Assessment Using Large Language Models.International Journal of Artificial Intelligence in Education35, 2 (01 Jun 2025), 559–586. doi:10.1007/s40593-...
2025 doi
-
[50]
Stanislav Pozdniakov, Jonathan Brazil, Solmaz Abdi, Aneesha Bakharia, Shazia Sadiq, Dragan Gašević, Paul Denny, and Hassan Khosravi. 2024. Large language models meet user interfaces: The case of provisioning feedback.Computers and Education: Artificial Intelligence7 (2024), 10...
2024
-
[51]
2009.The coding manual for qualitative researchers.Sage Publications Ltd, Thousand Oaks, CA
Johnny Saldaña. 2009.The coding manual for qualitative researchers.Sage Publications Ltd, Thousand Oaks, CA. xi, 223–xi, 223 pages
2009
-
[52]
Michael Sedlmair, Miriah Meyer, and Tamara Munzner. 2012. Design Study Methodology: Reflections from the Trenches and the Stacks.IEEE Transactions on Visualization and Computer Graphics18, 12 (2012), 2431–2440. doi:10.1109/TVCG. 2012.213
2012 doi
-
[53]
Karim Shabani, Mohamad Khatib, and Saman Ebadi. 2010. Vygotsky’s zone of proximal development: Instructional implications and teachers’ professional development.English language teaching3, 4 (2010), 237–248
2010
-
[54]
Yuhong Shi, Kun Yu, Yifei Dong, and Fang Chen. 2026. Large language models in education: a systematic review of empirical applications, benefits, and challenges.Computers and Education: Artificial Intelligence10 (2026), 100529. doi:10.1016/j.caeai.2025.100529
2026
-
[55]
Aarohi Srivastava et al. 2023. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.Transactions on Machine Learning Research(2023). https://openreview.net/forum?id=uyTL5Bvosj
2023
-
[56]
Jacob Steiss, Tamara Tate, Steve Graham, Jazmin Cruz, Michael Hebert, Jiali Wang, Youngsun Moon, Waverly Tseng, Mark Warschauer, and Carol Booth Olson. 2024. Comparing the quality of human and ChatGPT feedback of students’ writing.Learning and Instruction91 (2024), 101894. doi...
2024
-
[57]
Hendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover, Johanna Beyer, Hanspeter Pfister, and Alexander M. Rush. 2023. Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models. IEEE Transactions on Visualization and Computer Graphi...
2023
-
[58]
Karan Taneja, Pratyusha Maiti, Sandeep Kakar, Pranav Guruprasad, Sanjeev Rao, and Ashok K. Goel. 2024. Jill Watson: A Virtual Teaching Assistant Powered by ChatGPT. InArtificial Intelligence in Education, Andrew M. Olney, Irene-Angelica Chounta, Zitao Liu, Olga C. Santos, and ...
2024
-
[59]
Ehsan Toreini, Mhairi Aitken, Kovila Coopamootoo, Karen Elliott, Carlos Gonzalez Zelaya, and Aad van Moorsel
-
[60]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971(2023)
2023 arXiv
-
[61]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. InAdvances in Neural Information Processing Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vis...
2017
-
[62]
Qiaosi Wang, Koustuv Saha, Eric Gregori, David Joyner, and Ashok Goel. 2021. Towards Mutual Theory of Mind in Human-AI Interaction: How Language Reflects What Students Perceive About a Virtual Teaching Assistant. In Proceedings of the 2021 CHI Conference on Human Factors in Co...
2021
-
[63]
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuzni...
2022
-
[64]
Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Thomas Adams, Jeremy Howard, and Iacopo Poli. 2025. Smarter, Better, Faster, Longer: A Modern Bidirectional Enc...
2025
-
[65]
Sandra Wiktor, Aileen Benedict, and Mohsen Dorodchi. 2026. Supporting Peer-to-Peer Learning with LLMs: Investi- gating Smarter Student Solution Recommendations. InProceedings of the 57th ACM Technical Symposium on Computer Science Education V.1(USA)(SIGCSE TS 2026). Associatio...
2026
-
[66]
Magdalena Wischnewski, Nicole Krämer, and Emmanuel Müller. 2023. Measuring and Understanding Trust Calibrations for Automated Systems: A Survey of the State-Of-The-Art and Future Directions. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems(Hamburg...
2023
-
[67]
Do I Trust the AI?
Yuansong Xu, Yichao Zhu, Haokai Wang, Yuchen Wu, Yang Ouyang, Hanlu Li, Wenzhe Zhou, Xinyu Liu, Chang Jiang, and Quan Li. 2026. " Do I Trust the AI?" Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Reasoning.arXiv preprint arXiv:2601.1...
2026
-
[68]
Lixiang Yan, Lele Sha, Linxuan Zhao, Yuheng Li, Roberto Martinez-Maldonado, Guanliang Chen, Xinyu Li, Yueqiao Jin, and Dragan Gašević. 2024. Practical and ethical challenges of large language models in education: A systematic scoping review.British Journal of Educational Techn...
2024 doi
-
[69]
Zamfirescu-Pereira, Laryn Qi, Björn Hartmann, John DeNero, and Narges Norouzi
J.D. Zamfirescu-Pereira, Laryn Qi, Björn Hartmann, John DeNero, and Narges Norouzi. 2025. 61A Bot Report: AI Assistants in CS1 Save Students Homework Time and Reduce Demands on Staff. (Now What?). InProceedings of the 56th ACM Technical Symposium on Computer Science Education ...
2025
-
[70]
Wenxuan Zhang, Yue Deng, Bing Liu, Sinno Pan, and Lidong Bing. 2024. Sentiment Analysis in the Era of Large Language Models: A Reality Check. InFindings of the Association for Computational Linguistics: NAACL 2024, Kevin Duh, Helena Gomez, and Steven Bethard (Eds.). Associatio...
2024 doi
-
[71]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. InAdvances in Neural Information Proces...
2023
-
[72]
Huawei Zhou, Shuanghong Shen, Yu Su, Yongchun Miao, Qi Liu, Linbo Zhu, Junyu Lu, and Zhenya Huang. 2026. LLM-EPSP: Large language model empowered early prediction of student performance.Information Processing & Management63, 1 (2026), 104351. doi:10.1016/j.ipm.2025.104351 J. A...
2026
-
[2020]
InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency(Barcelona, Spain)(FAT* ’20)
The relationship between trust in AI and trustworthy machine learning technologies. InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency(Barcelona, Spain)(FAT* ’20). Association for Computing Machinery, New York, NY, USA, 272–283. doi:10.1145/3351...
2020
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.