Pith. sign in

REVIEW 4 major objections 4 minor 72 references

Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper claims that making trustworthiness visible—as five metrics and traceable visualizations—raises expert agreement when learning engineers evaluate LLM responses, and argues for hybrid human-in-the-loop evaluation.

desk verdict A well-executed co-design study with a plausible but under-supported central claim: the violation pipeline is unvalidated, so the observed IRR increase could be anchoring on shared noise. read the letter →

arxiv 2608.04006 v1 pith:NZ6UC77Q submitted 2026-08-04 cs.HC

classification cs.HC
keywords LLMevaluationtrustworthinesseducationtechnologyco-designvisualizationprompttournamentinter-raterreliabilitylearningengineers
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that making trustworthiness visible changes how learning engineers evaluate LLM responses in educational technology, and in the direction of better agreement. Through a year-long co-design with engineers building an LLM-powered digital textbook, the authors produced five trustworthiness metrics (20 measures) and visualizations that trace each violation to the sentence that caused it. In a prompt tournament with 12 learning engineers rating 30 real learner-sourced A/B match-ups, expert agreement rose when metrics were displayed, and engineers reported that the metrics acted as a checklist, surfacing pedagogical risks they had overlooked and making trade-offs between objectives explicit. The paper argues this supports hybrid evaluation: automated metrics screen for clear violations while humans resolve ambiguous, context-dependent cases. If right, trustworthiness is not just a benchmark property of a model but a workable decision-support lens for domain experts.

What carries the argument

The load-bearing mechanism is a per-response violation pipeline built on claim extraction and textual entailment. Each response's claims are checked against the textbook passage, the learner's summary, the chat history, and the latest learner message; entailment labels of entails/contradicts/neutral determine whether a claim counts as misinformation, hallucination, sycophancy, and so on—with neutral treated as the signature of an unsupported claim. Learning-method measures are scored by an LLM-as-a-judge against explicit pedagogical principles, and response-quality measures come from validated conversational features. The visualizations then map each flagged span back into the response text,

What would settle it

Run the same tournament with the same interfaces but with metric flags computed on shuffled or randomly mislabeled responses. If expert agreement rises as much with meaningless flags, the effect comes from having any shared display, not from trustworthy signal. Also, compare the automated flags to expert human judgments on the same 30 responses: near-chance agreement would directly undermine the claim.

Watch

Extended reading notes

Core claim

The central claim is that operationalizing LLM trustworthiness as a set of violation metrics plus text-traceable visualizations improves the reliability of expert evaluation of LLM responses in education. The authors co-constructed five metrics—Truthfulness, Diplomacy, Disposition, Learning Methods, and Response Quality—composed of 20 measures, each computed per response by an automated pipeline. They then ran an LLM prompt tournament in which 12 learning engineers chose between paired responses, half the time with the metrics and visualizations visible. Expert agreement was higher with metrics visible (alpha 0.4931 vs 0.3987), and the distribution of 'best' responses shifted toward lower-vi

Load-bearing premise

The whole finding rests on the assumption that the automated flags—built from claim extraction, textual entailment, an LLM judge, and fitted thresholds—are accurate enough to be trusted as evidence; the paper states their accuracy in educational contexts is not yet tested.

Editorial extensions

If this is right

  • If the claim holds, prompt-evaluation panels in education can raise agreement without imposing a single weighting of criteria.
  • Automated violation flags can serve as a first-pass screening layer, with engineers resolving ambiguous cases.
  • The metric set offers a reusable starting point for education-specific LLM evaluation that can be adjusted to other pedagogical contexts.
  • Comparison-oriented visualizations (summary overview, glyphs, donuts) are the features engineers actually use; single-metric encodings like opacity and underlining need refinement.
  • Evaluation tools that show trustworthiness alongside rubric criteria can change which prompt wins, since the rubric winner may be the least trustworthy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The agreement gain may come less from new information than from a shared reference point: any consistent, inspectable flag set might produce part of the effect, and a condition with plain-text metric summaries would isolate what visualization adds.
  • The same co-designed metric set could plausibly be adapted to other high-stakes LLM domains where experts disagree on outputs, since the mechanism is general rather than education-specific.
  • Because the study used one model and one dialogue task, the durability of the effect under repeated use and across models is open; a longitudinal replication would clarify whether trust in the measures grows or decays.
  • The dataset-derived thresholds (e.g., Participation, Diversity) should be recalibrated per corpus; without that, transferred deployments risk spurious flags that could erode the agreement benefit.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper reports a three-phase co-design study with learning engineers building an LLM-powered digital textbook. The authors first develop five trustworthiness metrics (Truthfulness, Diplomacy, Disposition, Learning Methods, Response Quality) comprising 20 measures, then design visualizations that map metric violations onto LLM responses, and finally evaluate the tools in an A/B prompt-evaluation tournament with 12 learning engineers rating 30 match-ups. The headline result is that making trustworthiness metrics visible increased inter-rater reliability (Krippendorff's alpha from 0.3987 without metrics to 0.4931 with metrics) and helped participants reconcile conflicting rubric objectives. The paper also contributes design guidelines for future LLM evaluation tools.

Significance. If the headline effect holds, this is a useful contribution to HCI and learning-engineering practice: it provides a concrete, domain-grounded operationalization of LLM trustworthiness for educational evaluation and a careful design-study template. The study has real strengths: it is ecologically valid (real student conversations, authentic prompt templates, within-subjects alternating interface conditions), it is longitudinal and co-design-based, and the authors are explicit about several key limitations—notably that the automated measures' accuracy in educational contexts is untested (§8.3) and that Participation and Diversity thresholds were fitted to the study dataset (§5.3). The qualitative findings—metrics acting as a procedural checklist, visualizations supporting trade-off reasoning—are credible and well supported by participant quotes. However, the central quantitative claim currently rests on an unvalidated measurement pipeline and an unsupported statistical comparison, which are the main barriers to accepting the paper's conclusions.

major comments (4)
  1. [§5.3, §8.3] The central claim presupposes that the 20 measures produce valid violation flags that learning engineers can use as trustworthy evidence. Section 8.3 states that 'their accuracy applied to educational contexts remains untested,' and no precision/recall, expert-label agreement, or any other validation is reported for any measure. The observed IRR increase (α0=0.3987 to α1=0.4931) is therefore also consistent with raters anchoring on a shared but possibly arbitrary overlay rather than on genuine trustworthiness information. Please add a validation study—e.g., expert labels on a held-out sample of responses with per-measure agreement, or at minimum a sensitivity analysis showing that the IRR effect survives when flagged-but-questionable responses are removed.
  2. [§7.1] The primary quantitative claim is the difference between α0=0.3987 and α1=0.4931. No confidence interval, bootstrap, or significance test is reported for either alpha or their difference, despite §4.3.6 adopting bootstrapped 95% CIs and the 'CIs do not overlap' rule. Both values are also below the conventional 0.67 reliability threshold, so the increase is from 'poor/moderate' to 'still poor/moderate.' Please report bootstrap CIs for each condition and a proper test of the difference (e.g., bootstrap percentile or permutation test); if the CIs overlap, the headline claim should be weakened accordingly.
  3. [§5.3 (Fig. 10Q,R)] The Participation and Diversity violation thresholds were, in the authors' own words, 'set from the distribution of our own dataset rather than an external standard' and 'require recalibration before being applied to another corpus.' Because the same dataset is used in the tournament, the 'With Metrics' condition displays flags that are partly fitted to the evaluation set. This is a circularity: the increase in agreement is not yet evidence for trustworthiness metrics in general, only for this dataset-calibrated instantiation. Please report thresholds derived from prior work or via cross-validation, and test the robustness of the IRR difference to reasonable threshold variations.
  4. [§4.3.3] Four of the twelve tournament raters participated in the earlier co-design phases for the metrics and visualizations (and, per §4.3.3, the additional eight were recruited later). Although participants had not seen the final visualizations, the co-designers helped define the design goals and measures; their prior investment may create demand characteristics or heightened fluency with the tool that the newly recruited raters lack. Please report whether the IRR increase (α0 vs. α1) holds when restricting the analysis to the eight non-co-design participants, or explicitly discuss this confound in the limitations.
minor comments (4)
  1. [§7.2 / Fig. 16] The time-on-task comparison is reported descriptively (median 79.55s vs. 62.95s; 6/12 slower). This is fine as an exploratory result, but it would fit the paper's own §4.3.6 statistical framework to report bootstrap CIs for the medians or means.
  2. [Fig. 17 / Fig. 19] The 'Ratio of (higher−lower)-violation responses picked' values are shown with CIs in Fig. 17, but Fig. 19 is presented without CIs or raw counts. Adding counts would help readers assess the strength of the 'selected in spite of violation' patterns.
  3. [Fig. 10O] The Dale-Chall readability formula and its constants are given only inside the figure. Since this is one of the 20 measures, please also define it in the text or a numbered equation so readers can reproduce the computation.
  4. [General] The CCS Concepts line still contains 'Do Not Use This Code' placeholders; this should be corrected before publication.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity; the central IRR finding is an empirical comparison, not a derivation, though dataset-fitted thresholds and an author-overlapping citation warrant noting.

full rationale

The paper's central claim—that making trustworthiness visible increased inter-rater reliability—is supported by a within-subjects prompt tournament experiment, not by a derivation from the paper's own equations. The disclosed fitting of Participation and Diversity thresholds 'from the distribution of our own dataset' (Section 5.3) is a validity limitation: the flags shown to raters are partly dataset-specific, but the outcome measure (IRR) is an empirical behavioral result, not a quantity forced by those thresholds. The paper explicitly states that 'their accuracy applied to educational contexts remains untested' (Section 8.3), which is an honest caveat about the pipeline rather than a circular reduction. The prompt tournament method is cited to the authors' prior work [27], but the method is fully described in Sections 3.1.2 and 4.3, so the citation is not load-bearing. The co-designers later serving as evaluators creates a reflexivity concern, but the IRR metric is computed from blind A/B choices and the paper acknowledges such individual-difference factors. Under the hard rule requiring an explicit reduction (e.g., Eq. X = Eq. Y by construction, or a fitted parameter renamed as a prediction), no such circular step can be quoted from this paper.

Assumptions & free parameters 7 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the accuracy of the violation pipeline, the validity of LLM-as-a-judge outputs, and the representativeness of a small co-designed study. Two thresholds are fitted to the evaluation dataset itself, which the authors disclose. No new physical or conceptual entities are introduced beyond the metrics and measures, which are constructs rather than entities in the graviton sense.

free parameters (7)
  • Entailment confidence threshold = 50%
    An entailment label is accepted only above 50% confidence; all truthfulness and diplomacy violations depend on this threshold.
  • Semantic similarity threshold = 0.5 (text) / 0.7 (figure example)
    Two claims are treated as semantically similar above cosine 0.5 in the implementation, while the figure uses 0.7 as an example; affects sycophancy, adversarial factuality, emotional awareness, and preference bias.
  • Certainty violation threshold = < 5.5
    Responses scoring below 5.5 on the Certainty Lexicon are flagged; the threshold is chosen by the research team.
  • Readability violation threshold = Dale-Chall score >= 0.6
    Responses above this Dale-Chall score are flagged as difficult to read.
  • Subjectivity violation threshold = > 0.5
    Responses with a personal-to-factual ratio above 0.5 are flagged as subjective.
  • Participation violation threshold = Gini > 0.4
    The paper states this threshold was set from the distribution of the authors' own dataset rather than an external standard.
  • Diversity violation threshold = topic vector score > 0.7
    The paper states this threshold was set from the distribution of the authors' own dataset rather than an external standard.
assumptions (5)
  • domain assumption The automated violation pipeline produces flags accurate enough to support expert decisions
    The authors explicitly say in Section 8.3 that the accuracy of their measures applied to educational contexts remains untested; if flags are noisy, the observed decision effects may reflect noise.
  • domain assumption Textual entailment and semantic similarity models transfer to learner-generated text and textbook passages
    Claim extraction and entailment are run on real student summaries and chatbot dialogue, but no validation is provided for this genre.
  • domain assumption LLM-as-a-Judge returns valid violations for Socratic, Scaffolding, Connections, Metacognitive, and SERT
    Five learning-method measures rely entirely on gpt-5 judging whether a response violates four principles per method; this delegation is cited as common practice but is not validated here.
  • domain assumption The 12 learning engineer raters are representative of the broader learning engineer population
    The study relies on a small convenience sample of collaborators and additional recruited engineers; generalizability to other EdTech teams is assumed.
  • domain assumption The alternating within-subjects interface design controls for learning and ordering effects
    Practice tasks only use the with-metrics interface, and carryover of metric awareness into without-metrics trials cannot be fully excluded.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education." pith.science (2026). https://pith.science/paper/NZ6UC77Q

@misc{pith2026260804006,
  author       = {Pith},
  title        = {Pith review of: Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NZ6UC77Q}},
  note         = {Machine review of arXiv:2608.04006}
}
read the original abstract

LLMs are reshaping educational technology, yet evaluating their responses for pedagogical alignment remains underexplored, relying heavily on the expertise of learning engineers building the technology. To bridge this gap, we explore trustworthiness as a structured lens for evaluation, leveraging existing measures of LLM trustworthiness to systematically identify potential pedagogical disruptions. Through a longitudinal co-design process with learning engineers developing an LLM-powered digital textbook, we: (1) co-constructed five trustworthiness metrics comprising 20 measures tailored to pedagogical use; (2) designed visualizations that map trustworthiness violations onto LLM responses; and (3) evaluated how these tools help learning engineers make A/B comparisons of LLM responses. Making trustworthiness explicit increased inter-rater reliability while helping learning engineers resolve conflicting objectives and produce more consistent judgments. We discuss the emergent benefits of trustworthiness as a lens for evaluating LLMs in education and propose new design guidelines for future evaluation tools that enable pedagogically-aligned, LLM-powered learning tools.

Figures

Figures reproduced from arXiv: 2608.04006 by the authors.

Figure 1
Figure 1. To help learning engineers calibrate their evaluations of LLMs in education, we [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Timeline of our collaborative co-design process with learning engineers. We continuously discussed [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Our initial collection of 3 metric categories, composed of 27 individual measures (18 trust-based, 5 [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (19 more)
Figure 4
Figure 4. Figure 4: The initial designs for visualizations of our trustworthiness metrics. We used a whiteboarding-style [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: An overview of our LLM prompt tournament dataset. We collected 30 inputs sourced from real student [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: The tournament interfaces we developed as our two study conditions. The “Without Metrics” interface [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: The rubric for the prompt tournament, describing the study task for each rater. The rubric was [PITH_FULL_IMAGE:figures/full_fig_p019_7.png]
Figure 8
Figure 8. Figure 8: An overview of our study design for conducting an LLM prompt evaluation tournament. [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: An overview of our final 5 trustworthiness metrics, composed of 20 individual measures. Our metrics [PITH_FULL_IMAGE:figures/full_fig_p023_9.png]
Figure 10
Figure 10. Figure 10: A breakdown of how we calculate violations of all 20 measures. [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]
Figure 11
Figure 11. Figure 11: An example of calculating violations of Hallucination and Sycophancy measures on a sample input. [PITH_FULL_IMAGE:figures/full_fig_p025_11.png]
Figure 12
Figure 12. Figure 12: An overview of our final metrics visualizations. [PITH_FULL_IMAGE:figures/full_fig_p028_12.png]
Figure 13
Figure 13. Figure 13: To visualize potentially overlapping measure violations in an LLM response across multiple metrics [PITH_FULL_IMAGE:figures/full_fig_p029_13.png]
Figure 14
Figure 14. Figure 14: Raw win rates for each LLM prompt template across both conditions (left) and broken down by [PITH_FULL_IMAGE:figures/full_fig_p031_14.png]
Figure 15
Figure 15. Figure 15: Count of measure violations in all responses per LLM. Prompt 3, the winner of the prompt tournament, [PITH_FULL_IMAGE:figures/full_fig_p032_15.png]
Figure 16
Figure 16. Figure 16: Time taken by the participants to choose a response for each interface condition. Overall, having [PITH_FULL_IMAGE:figures/full_fig_p034_16.png]
Figure 17
Figure 17. Figure 17: The ratio of (higher - lower) -violation responses chosen over total per participant across interface [PITH_FULL_IMAGE:figures/full_fig_p035_17.png]
Figure 18
Figure 18. Figure 18: 95% CI around mean self-reported frequency ([1, 5]) and usefulness ([−3, 3]) of each measure. Misin￾formation, Hallucination, Feedback, Socratic, Scaffolding, and Readability were both frequently used and useful. Adversarial Factuality, Preference Bias, Connections, M…
Figure 19
Figure 19. Figure 19: The percentage of measure violations in all responses chosen per LLM. Higher values and lighter [PITH_FULL_IMAGE:figures/full_fig_p037_19.png]
Figure 20
Figure 20. Figure 20: Table of observed participant workflows during the prompt tournament. [PITH_FULL_IMAGE:figures/full_fig_p038_20.png]
Figure 21
Figure 21. Figure 21: 95% CI around mean self-reported frequency ([1, 5]), usability ([−3, 3]), and usefulness ([−3, 3]) ratings of interface features. The glyphs and explanations were used the most often and considered highly usable and useful. Both the measure and metric text highlightin…
Figure 22
Figure 22. Figure 22: We also developed a “Text Characteristics” panel that was co-designed with our collaborators, but [PITH_FULL_IMAGE:figures/full_fig_p043_22.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

72 extracted references · 28 canonical work pages

  1. [1]

    Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent Anti-Muslim Bias in Large Language Models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society(Virtual Event, USA)(AIES ’21). Association for Computing Machinery, New York, NY, USA, 298–306. doi:10.1145/3461702.3462624

  2. [2]

    Saleh Afroogh, Ali Akbari, Emmie Malone, Mohammadali Kargar, and Hananeh Alambeigi. 2024. Trust in AI: progress, challenges, and future directions.Humanities and Social Sciences Communications11, 1 (2024), 1568

  3. [3]

    Boyatzis

    R.E. Boyatzis. 1998.Transforming Qualitative Information: Thematic Analysis and Code Development. SAGE Publications

  4. [4]

    Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, et al. 2020. Toward trustworthy AI development: mechanisms for supporting verifiable claims.arXiv preprint arXiv:2004.07213(2020)

  5. [5]

    Davide Ceneda, Theresia Gschwandtner, Thorsten May, Silvia Miksch, Hans-Jörg Schulz, Marc Streit, and Christian Tominski. 2017. Characterizing Guidance in Visual Analytics.IEEE Transactions on Visualization and Computer Graphics 23, 1 (2017), 111–120. doi:10.1109/TVCG.2016.2598468

  6. [6]

    Chall and Edgar Dale

    Jeanne S. Chall and Edgar Dale. 1995.Readability Revisited: The New Dale-Chall Readability Formula. Brookline Books, Cambridge, MA, USA

  7. [7]

    Angelos Chatzimparmpas, Kostiantyn Kucher, and Andreas Kerren. 2024. Visualization for Trust in Machine Learning Revisited: The State of the Field in 2023.IEEE Computer Graphics and Applications44, 3 (2024), 99–113. doi:10.1109/ MCG.2024.3360881

  8. [8]

    Chatzimparmpas, R

    A. Chatzimparmpas, R. M. Martins, I. Jusufi, K. Kucher, F. Rossi, and A. Kerren. 2020. The State of the Art in Enhancing Trust in Machine Learning Models with the Use of Visualizations.Computer Graphics Forum39, 3 (2020), 713–756. doi:10.1111/cgf.14034

Show all 72 references
  1. [9]

    Lijia Chen, Pingping Chen, and Zhijian Lin. 2020. Artificial Intelligence in Education: A Review.IEEE Access8 (2020), 75264–75278. doi:10.1109/ACCESS.2020.2988510

  2. [10]

    Adam Coscia and Alex Endert. 2024. KnowledgeVIS: Interpreting Language Models by Comparing Fill-in-the-Blank Prompts.IEEE Transactions on Visualization and Computer Graphics30, 9 (2024), 6520–6532. doi:10.1109/TVCG.2023. 3346713

  3. [11]

    Choi, Scott Crossley, and Alex Endert

    Adam Coscia, Langdon Holmes, Wesley Morris, Joon S. Choi, Scott Crossley, and Alex Endert. 2024. iScore: Visual Analytics for Interpreting How Language Models Automatically Score Summaries. InProceedings of the 2024 IUI Conference on Intelligent User Interfaces(Greenville, SC,...

  4. [12]

    Adam J Coscia, Shunan Guo, Eunyee Koh, and Alex Endert. 2025. OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST ’25). Associati...

  5. [13]

    Cotton Debby R

    Peter A. Cotton Debby R. E. Cotton and J. Reuben Shipway. 2024. Chatting and cheating: Ensuring academic integrity in the era of ChatGPT.Innovations in Education and Teaching International61, 2 (2024), 228–239. arXiv:https://doi.org/10.1080/14703297.2023.2190148 doi:10.1080/14...

  6. [14]

    2018.Learning engineering for online education

    Chris Dede, Bror Saxberg, and John Richards. 2018.Learning engineering for online education. Vol. 95. Routledge New York

  7. [15]

    Andre del Carpio Gutierrez, Paul Denny, and Andrew Luxton-Reilly. 2024. Automating Personalized Parsons Problems with Customized Contexts and Concepts(ITiCSE 2024). Association for Computing Machinery, New York, NY, USA, 688–694. doi:10.1145/3649217.3653568

  8. [16]

    2016.Fair Statistical Communication in HCI

    Pierre Dragicevic. 2016.Fair Statistical Communication in HCI. Springer International Publishing, Cham, 291–330. doi:10.1007/978-3-319-26633-6_13

  9. [17]

    Jessica Díaz, Jorge Pérez, Carolina Gallardo, and Ángel González-Prieto. 2023. Applying Inter-Rater Reliability and Agreement in collaborative Grounded Theory studies in software engineering.Journal of Systems and Software195 (2023), 111520. doi:10.1016/j.jss.2022.111520

  10. [18]

    Evans and William Revelle

    Anthony M. Evans and William Revelle. 2008. Survey and behavioral measurements of interpersonal trust.Journal of Research in Personality42, 6 (2008), 1585–1593. doi:10.1016/j.jrp.2008.07.011

  11. [19]

    Niles, Ken Pathak, and Steven Sloan

    Md Meftahul Ferdaus, Mahdi Abdelguerfi, Elias Loup, Kendall N. Niles, Ken Pathak, and Steven Sloan. 2026. Towards Trustworthy AI: A Review of Ethical and Robust Large Language Models.ACM Comput. Surv.58, 7, Article 176 (Jan. 2026), 43 pages. doi:10.1145/3777382

  12. [20]

    Yingqiang Ge, Shuchang Liu, Zuohui Fu, Juntao Tan, Zelong Li, Shuyuan Xu, Yunqi Li, Yikun Xian, and Yongfeng Zhang. 2024. A Survey on Trustworthy Recommender Systems.ACM Trans. Recomm. Syst.3, 2, Article 13 (Nov. 2024), 68 pages. doi:10.1145/3652891

  13. [21]

    Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2020, Trevor Cohn, Yulan He, and Yang Liu (Eds.)...

  14. [22]

    Kummerfeld, and Elena L

    Katy Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K. Kummerfeld, and Elena L. Glassman. 2024. Supporting Sensemaking of Large Language Model Outputs at Scale. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Associa...

  15. [23]

    Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A Survey on Automated Fact-Checking.Transactions of the Association for Computational Linguistics10 (2022), 178–206. doi:10.1162/tacl_a_00454

  16. [24]

    Silva, and Huamin Qu

    Jianben He, Xingbo Wang, Shiyi Liu, Guande Wu, Cláudio T. Silva, and Huamin Qu. 2025. POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models. In2025 IEEE 18th Pacific Visualization Conference (PacificVis). 36–46. doi:10.1109/PACIFICVI...

  17. [25]

    Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. InProceedings of the 9th International Conference on Learning Representations (ICLR ’21). arXiv:2006.03654

  18. [26]

    Fred Hohman, Minsuk Kahng, Robert Pienta, and Duen Horng Chau. 2019. Visual Analytics in Deep Learning: An Interrogative Survey for the Next Frontiers.IEEE Transactions on Visualization and Computer Graphics25, 8 (2019), 2674–2693. doi:10.1109/TVCG.2018.2843369

  19. [27]

    Langdon Holmes, Adam Coscia, Scott Crossley, Joon Suh Choi, and Wesley Morris. 2026. LLM Prompt Evaluation for Educational Applications.arXiv preprint arXiv:2601.16134(2026)

  20. [28]

    Santos, Mercedes T

    Wayne Holmes, Kaska Porayska-Pomsta, Ken Holstein, Emma Sutherland, Toby Baker, Simon Buckingham Shum, Olga C. Santos, Mercedes T. Rodrigo, Mutlu Cukurova, Ig Ibert Bittencourt, and Kenneth R. Koedinger. 2022. Ethics of AI in Education: Towards a Community-Wide Framework.Inter...

  21. [29]

    Xinlan Emily Hu. 2025. team_comm_tools: A Python Toolkit for Exploring the Communication Space. PsyArXiv preprint. doi:10.31234/osf.io/pz98q_v2

  22. [30]

    Yue Huang et al. 2024. Position: TrustLLM: Trustworthiness in Large Language Models. InForty-first International Conference on Machine Learning. https://openreview.net/forum?id=bWUU0LwwMp

  23. [31]

    Bill Hunter. 2015. Teaching for engagement: Part 1–Constructivist principles, case-based teaching, and active learning. College Quarterly18, 2 (2015), n2. J. ACM, Vol. 37, No. 4, Article 111. Publication date: August 2026. 111:46 Coscia et al

  24. [32]

    Muhammad Imran and Norah Almusharraf. 2024. Google Gemini as a next generation AI educational tool: a review of emerging educational technology.Smart Learning Environments11, 1 (23 May 2024), 22. doi:10.1186/s40561-024-00310-z

  25. [33]

    Ellen Jiang, Kristen Olson, Edwin Toh, Alejandra Molina, Aaron Donsbach, Michael Terry, and Carrie J. Cai. 2022. PromptMaker: Prompt-based Prototyping with Large Language Models. InExtended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (CHI EA ’22)...

  26. [34]

    Minsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu, James Wexler, Emily Reif, Krystal Kallarackal, Minsuk Chang, Michael Terry, and Lucas Dixon. 2024. LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models. InExtended Abstracts of th...

  27. [35]

    Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al. 2023. ChatGPT for good? On opportunities and challenges of large language models for education.Learning and...

  28. [36]

    Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2024. EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Associat...

  29. [37]

    Malcolm S. Knowles. 1978. Andragogy: Adult Learning Theory in Perspective.Community College Review5, 3 (1978), 9–20. doi:10.1177/009155217800500302

  30. [38]

    Katerina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé, Danai Myrtzani, Theodoros Evgeniou, Ion Androut- sopoulos, and John Pavlopoulos. 2025. Evaluation and Facilitation of Online Discussions in the LLM Era: A Survey. In Proceedings of the 2025 Conference on Empirical ...

  31. [39]

    Moritz Laurer, Wouter van Atteveldt, Andreu Casas, and Kasper Welbers. 2024. Less Annotating, More Classifying: Addressing the Data Scarcity Issue of Supervised Machine Learning with Deep Transfer Learning and BERT-NLI. Political Analysis32, 1 (2024), 84–100. doi:10.1017/pan.2023.20

  32. [40]

    Maxime Lelièvre, Amy Waldock, Meng Liu, Natalia Valdés Aspillaga, Alasdair Mackintosh, María José Ogando Portela, Jared Lee, Paul Atherton, Robin AA Ince, and Oliver GB Garrod. 2025. Benchmarking the pedagogical knowledge of large language models.arXiv preprint arXiv:2506.18710(2025)

  33. [41]

    Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou. 2023. Trustworthy AI: From Principles to Practices.ACM Comput. Surv.55, 9, Article 177 (Jan. 2023), 46 pages. doi:10.1145/3555803

  34. [42]

    Jakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur, Iryna Gurevych, and Mrinmaya Sachan. 2025. MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors.arXiv preprint arXiv:2502.18940 (2025)

  35. [43]

    McNamara

    Danielle S. McNamara. 2004. SERT: Self-Explanation Reading Training.Discourse Processes38, 1 (2004), 1–30. doi:10.1207/s15326950dp3801_1

  36. [44]

    Meyer, Ryan J

    Jesse G. Meyer, Ryan J. Urbanowicz, Patrick C. N. Martin, Karen O’Connor, Ruowang Li, Pei-Chen Peng, Tiffani J. Bright, Nicholas Tatonetti, Kyoung Jae Won, Graciela Gonzalez-Hernandez, and Jason H. Moore. 2023. ChatGPT and large language models in academia: opportunities and c...

  37. [45]

    Mihran Miroyan, Chancharik Mitra, Rishi Jain, Gireeja Ranade, and Narges Norouzi. 2025. Analyzing Pedagogical Quality and Efficiency of LLM Responses with TA Feedback to Live Student Questions. InProceedings of the 56th ACM Technical Symposium on Computer Science Education V. ...

  38. [47]

    Wesley Morris, Scott Crossley, Langdon Holmes, Chaohua Ou, Mihai Dascalu, and Danielle McNamara. 2025. Formative Feedback on Student-Authored Summaries in Intelligent Textbooks Using Large Language Models.International Journal of Artificial Intelligence in Education35, 3 (01 S...

  39. [48]

    Wesley Morris, Scott Crossley, Langdon Holmes, Chaohua Ou, Danielle McNamara, and Mihai Dascalu. 2023. Using Large Language Models to Provide Formative Feedback in Intelligent Textbooks. InArtificial Intelligence in Education. Posters and Late Breaking Results, Workshops and T...

  40. [49]

    Wesley Morris, Langdon Holmes, Joon Suh Choi, and Scott Crossley. 2025. Automated Scoring of Constructed Response Items in Math Assessment Using Large Language Models.International Journal of Artificial Intelligence in Education35, 2 (01 Jun 2025), 559–586. doi:10.1007/s40593-...

  41. [50]

    Stanislav Pozdniakov, Jonathan Brazil, Solmaz Abdi, Aneesha Bakharia, Shazia Sadiq, Dragan Gašević, Paul Denny, and Hassan Khosravi. 2024. Large language models meet user interfaces: The case of provisioning feedback.Computers and Education: Artificial Intelligence7 (2024), 10...

  42. [51]

    2009.The coding manual for qualitative researchers.Sage Publications Ltd, Thousand Oaks, CA

    Johnny Saldaña. 2009.The coding manual for qualitative researchers.Sage Publications Ltd, Thousand Oaks, CA. xi, 223–xi, 223 pages

  43. [52]

    Michael Sedlmair, Miriah Meyer, and Tamara Munzner. 2012. Design Study Methodology: Reflections from the Trenches and the Stacks.IEEE Transactions on Visualization and Computer Graphics18, 12 (2012), 2431–2440. doi:10.1109/TVCG. 2012.213

  44. [53]

    Karim Shabani, Mohamad Khatib, and Saman Ebadi. 2010. Vygotsky’s zone of proximal development: Instructional implications and teachers’ professional development.English language teaching3, 4 (2010), 237–248

  45. [54]

    Yuhong Shi, Kun Yu, Yifei Dong, and Fang Chen. 2026. Large language models in education: a systematic review of empirical applications, benefits, and challenges.Computers and Education: Artificial Intelligence10 (2026), 100529. doi:10.1016/j.caeai.2025.100529

  46. [55]

    Aarohi Srivastava et al. 2023. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.Transactions on Machine Learning Research(2023). https://openreview.net/forum?id=uyTL5Bvosj

  47. [56]

    Jacob Steiss, Tamara Tate, Steve Graham, Jazmin Cruz, Michael Hebert, Jiali Wang, Youngsun Moon, Waverly Tseng, Mark Warschauer, and Carol Booth Olson. 2024. Comparing the quality of human and ChatGPT feedback of students’ writing.Learning and Instruction91 (2024), 101894. doi...

  48. [57]

    Hendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover, Johanna Beyer, Hanspeter Pfister, and Alexander M. Rush. 2023. Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models. IEEE Transactions on Visualization and Computer Graphi...

  49. [58]

    Karan Taneja, Pratyusha Maiti, Sandeep Kakar, Pranav Guruprasad, Sanjeev Rao, and Ashok K. Goel. 2024. Jill Watson: A Virtual Teaching Assistant Powered by ChatGPT. InArtificial Intelligence in Education, Andrew M. Olney, Irene-Angelica Chounta, Zitao Liu, Olga C. Santos, and ...

  50. [59]

    Ehsan Toreini, Mhairi Aitken, Kovila Coopamootoo, Karen Elliott, Carlos Gonzalez Zelaya, and Aad van Moorsel

  51. [60]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971(2023)

  52. [61]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. InAdvances in Neural Information Processing Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vis...

  53. [62]

    Qiaosi Wang, Koustuv Saha, Eric Gregori, David Joyner, and Ashok Goel. 2021. Towards Mutual Theory of Mind in Human-AI Interaction: How Language Reflects What Students Perceive About a Virtual Teaching Assistant. In Proceedings of the 2021 CHI Conference on Human Factors in Co...

  54. [63]

    Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuzni...

  55. [64]

    Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Thomas Adams, Jeremy Howard, and Iacopo Poli. 2025. Smarter, Better, Faster, Longer: A Modern Bidirectional Enc...

  56. [65]

    Sandra Wiktor, Aileen Benedict, and Mohsen Dorodchi. 2026. Supporting Peer-to-Peer Learning with LLMs: Investi- gating Smarter Student Solution Recommendations. InProceedings of the 57th ACM Technical Symposium on Computer Science Education V.1(USA)(SIGCSE TS 2026). Associatio...

  57. [66]

    Magdalena Wischnewski, Nicole Krämer, and Emmanuel Müller. 2023. Measuring and Understanding Trust Calibrations for Automated Systems: A Survey of the State-Of-The-Art and Future Directions. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems(Hamburg...

  58. [67]

    Do I Trust the AI?

    Yuansong Xu, Yichao Zhu, Haokai Wang, Yuchen Wu, Yang Ouyang, Hanlu Li, Wenzhe Zhou, Xinyu Liu, Chang Jiang, and Quan Li. 2026. " Do I Trust the AI?" Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Reasoning.arXiv preprint arXiv:2601.1...

  59. [68]

    Lixiang Yan, Lele Sha, Linxuan Zhao, Yuheng Li, Roberto Martinez-Maldonado, Guanliang Chen, Xinyu Li, Yueqiao Jin, and Dragan Gašević. 2024. Practical and ethical challenges of large language models in education: A systematic scoping review.British Journal of Educational Techn...

  60. [69]

    Zamfirescu-Pereira, Laryn Qi, Björn Hartmann, John DeNero, and Narges Norouzi

    J.D. Zamfirescu-Pereira, Laryn Qi, Björn Hartmann, John DeNero, and Narges Norouzi. 2025. 61A Bot Report: AI Assistants in CS1 Save Students Homework Time and Reduce Demands on Staff. (Now What?). InProceedings of the 56th ACM Technical Symposium on Computer Science Education ...

  61. [70]

    Wenxuan Zhang, Yue Deng, Bing Liu, Sinno Pan, and Lidong Bing. 2024. Sentiment Analysis in the Era of Large Language Models: A Reality Check. InFindings of the Association for Computational Linguistics: NAACL 2024, Kevin Duh, Helena Gomez, and Steven Bethard (Eds.). Associatio...

  62. [71]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. InAdvances in Neural Information Proces...

  63. [72]

    Huawei Zhou, Shuanghong Shen, Yu Su, Yongchun Miao, Qi Liu, Linbo Zhu, Junyu Lu, and Zhenya Huang. 2026. LLM-EPSP: Large language model empowered early prediction of student performance.Information Processing & Management63, 1 (2026), 104351. doi:10.1016/j.ipm.2025.104351 J. A...

  64. [2020]

    InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency(Barcelona, Spain)(FAT* ’20)

    The relationship between trust in AI and trustworthy machine learning technologies. InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency(Barcelona, Spain)(FAT* ’20). Association for Computing Machinery, New York, NY, USA, 272–283. doi:10.1145/3351...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.