Pith. sign in

REVIEW 3 major objections 4 minor 87 references

PlanGlow: Personalized Study Planning with an Explainable and Controllable LLM-Driven System

T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read PlanGlow claims that adding explanations and user control to an LLM study planner beats a bare GPT-4o prompt box and Khanmigo on usability, explainability, controllability, and expert-rated plan quality.

desk verdict Solid, honest systems paper whose headline claims run ahead of a feature-unmatched baseline and single-rater expert scoring; the design patterns are still worth engaging with. read the letter →

arxiv 2504.12452 v1 pith:2TX5HP4I submitted 2025-04-16 cs.HC

classification cs.HC
keywords Self-directedlearningPersonalizedExplainableAIControllableLargelanguagemodelsStudyplanningUserhallucination
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

PlanGlow is an LLM-based study-planning system built to test a specific claim: that learners' problems with AI-generated study plans—opaque reasoning, hallucinated or mismatched resources, and difficulty adjusting the plan—can be addressed by design features rather than by a better model. The paper reports a within-subject study in which 24 learners compared PlanGlow with a bare GPT-4o prompt box and with Khanmigo, and one educator scored the resulting plans against criteria developed with a second educator. PlanGlow scored significantly higher on functional integration, efficient plan generation, reliable resource validation, and on explanations of recommendation rationale, goal alignment, and task connections, and it received significantly higher expert ratings on objectives, timelines, resources, and pedagogical soundness. Raw performance and most usability items showed no significant difference. If the claim is right, the takeaway for AI learning tools is that explanations and user control, not the underlying model, drive measurable gains in perceived plan quality at the planning stage of self-directed learning.

What carries the argument

The load-bearing object is PlanGlow's plan-generation and presentation pipeline: a structured input form collects subject, goals, background level, duration, and daily availability; a three-step chain-of-thought procedure (initial generation, critique, improvement) produces the plan; a layered interface exposes weekly overviews, daily breakdowns, rationale boxes, and video-validation status; and in-line editing, chat, and resource replacement give users control. The features that carry the argument are the explanation panels (rationale, objectives, connections) and the validated resource list, because they distinguish PlanGlow from chat-based baselines and map directly to the hypotheses that showed significant gains.

What would settle it

Run the same 24-participant within-subject comparison against a feature-matched control: the same structured form, editing, and resource list as PlanGlow, but with the rationale boxes, explanation toggles, and video-validation badges removed. If users no longer rate the control lower on H3b, H3d, and H4d–H4g, then the features are not the driver; if they do, the paper's interpretation holds. For the expert ratings, two independent raters scoring full plans in identical plain-text formatting would test whether the H5 advantage survives presentation bias.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that wrapping an LLM study-plan generator in explanation and control features moves user-perceived and expert-rated plan quality. In a within-subject comparison with a GPT-4o prompt box and Khanmigo, PlanGlow scored significantly higher on functional integration (H2d), efficient generation of the desired plan (H3b), and reliable resource validation (H3d), and, against Khanmigo, on easy plan generation (H3a) and straightforward alternative-resource search (H3c). It also scored significantly higher on explaining the rationale behind recommendations, aligning plans with goals, clarifying daily-weekly task connections, and enabling informed decisions (H4d–H4g), and on concise explanations versus Khanmigo (H4a). Expert ratings favored PlanGlow on learning objectives, timelines, resources, and pedagogical soundness against both baselines (H5a–H5c, H5e), while progress monitoring was not significantly better than GPT-4o (H5d). Overall performance, most usability items, and explanation accuracy or relevance did not differ significantly, and 83.3% of participants ranked PlanGlow as their top choice.

Load-bearing premise

The claim rests on the assumption that PlanGlow's measured advantages come from its explanation and control features rather than from simply having a purpose-built interface, since its main comparator was a bare text box and only one expert rated simplified plan excerpts.

Editorial extensions

If this is right

  • If the results hold, adding structured rationale and resource validation to an LLM planner is enough to move user-perceived plan quality even when the underlying model stays the same.
  • Learning platforms that already offer LLM chat could adopt the form-plus-editing-plus-explanation pattern without changing their model.
  • Because the usability win was functional integration rather than general ease of use, the design lesson is feature coherence, not overall polish.
  • Planning-stage tools for self-directed learners can expect expert-rated objectives, timelines, resources, and pedagogy to improve when plans carry explicit rationale, and should not rely on raw model output alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves open whether users rate explanations highly because they genuinely inform decisions or because they signal care and structure; a follow-up that measures decision quality, not just ratings, would separate the two.
  • Because the GPT-4o comparator was a bare text box, part of the gap may come from having any purpose-built interface; a feature-matched control that removes only the rationale panels would isolate the explainability effect.
  • The low observed use of in-line editing and chat suggests that the visible presence of control may matter more than actual manipulation in a single session, so whether control features pay off over weeks of self-directed learning is a natural longitudinal test.
  • The absence of significant gains on explanation accuracy or relevance and on progress monitoring implies PlanGlow shifts perceived quality and planning structure without yet demonstrating improved factual reliability or execution support; coupling the system with retrieval or assessment tools would test whether those gaps close.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper presents PlanGlow, a web-based LLM-driven study planning system with structured input forms, generated plans with explanations (rationale, learning objectives, task connections), YouTube API-based resource validation, in-line editing, and a chat feature. The authors conducted a formative survey (n=28) and interviews (n=10 plus one educational researcher) to derive four design requirements, then ran a within-subject experiment with 24 students comparing PlanGlow to a GPT-4o prompt-box baseline and Khan Academy's Khanmigo. Outcome measures were 7-point Likert ratings on performance, usability, controllability, and explainability (H1-H4) and expert ratings of plan quality (H5). Results show no significant performance differences (H1 rejected), a single usability win on functional integration (H2d), significant controllability gains on plan-generation ease vs Khanmigo, efficiency, resource search vs Khanmigo, and resource validation reliability (H3a-H3d), significant explainability gains on conciseness vs Khanmigo and on rationale, goal alignment, task connections, and informed decisions (H4d-H4g), and expert-rated quality advantages on multiple H5 sub-hypotheses. The abstract claims that PlanGlow 'significantly improves usability, explainability, and controllability,' and 83.3% of participants ranked PlanGlow as their top choice.

Significance. If the reported results are taken at face value, this is a useful empirical evaluation of an LLM-based tool for the planning stage of self-directed learning, an area the paper correctly identifies as underexplored. The manuscript has several strengths: a counterbalanced within-subject design, pre-specified hypotheses, Bonferroni post-hoc corrections, detailed interaction logs, honest reporting of rejected hypotheses (H1, most of H2, H3e, H4b, H4c, H5d), and publicly available code and evaluation materials. The system design, including chain-of-thought generation with a critique and improvement step and YouTube API-based resource validation, is concrete and reproducible. The main limitation is interpretive: the measured advantages are not cleanly attributable to the named explainability and controllability constructs because the comparators differ in many interface dimensions. The honest null results and high engagement with explanation features are nevertheless valuable for future XAI-in-education work.

major comments (3)
  1. [§5.2, Figure 4; §6, H2-H4 results] The attribution of PlanGlow's advantages to explainability and controllability is confounded by the feature-unmatched baseline. The GPT-4o comparator is a bare text box with no structured form, plan visualization, editing affordances, resource display, or explicit explanation blocks. Because H3d (resource validation reliability), H4d-H4g (rationale, goal alignment, task connections, informed decisions), and H2d (functional integration) are measured at the feature level, the significant differences could simply reflect the presence of any purpose-built interface rather than the quality of the explainability and control designs. To support the headline claim, the paper would need a feature-matched control or ablation (e.g., PlanGlow with explanations and controls disabled). As it stands, the evidence supports a system-level comparison ('a purpose-built plan UI was preferred to a prompt box'), not a causal attribution to the named constructs.
  2. [§5.5 and Table 3; abstract; §1] The expert evaluation protocol is not as described in the abstract. Section 5.5 states that E1 evaluated all generated plans, while E2 only collaborated on developing the evaluation criteria; the abstract and Introduction's claim that 'two educational experts assessed' the plans overstates the protocol. There is also no inter-rater reliability coefficient. In addition, PlanGlow plans were presented as simplified Week-1/Day-1 excerpts while E1 was told that full explanations existed and resources had been validated, whereas the GPT-4o and Khanmigo plans were presumably presented in full text. This differential presentation and expectation cue could plausibly bias the H5a-H5c and H5e ratings. Please report the exact presentation format for all three systems, justify the belief that the excerpts are representative, and either provide a second independent rater or clearly frame H5 as a single-rater exploratory assessment.
  3. [Abstract; §8 Conclusion] The abstract's claim that PlanGlow 'significantly improves usability, explainability, and controllability' is not supported by the reported statistics. Usability was significant only for H2d (functional integration); H2a-H2c and H2e-H2g were rejected. Explainability was significant for conciseness only against Khanmigo and was not significant for accuracy (H4b) or relevance (H4c). The Conclusion's statement that PlanGlow 'significantly outperformed two baselines' is likewise overly broad given the many non-significant comparisons. The abstract and conclusion should be revised to state precisely which sub-hypotheses were supported, or the paper should present a composite measure (if justified) before claiming overall improvements in those dimensions.
minor comments (4)
  1. [§4, Figure 1 caption and text] The text contains garbled notation and a duplicated phrase: 'the "bulb" ♂lightbulbicon' should be cleaned up, and 'the system provides a clear visual indicator provides real-time feedback' repeats the verb 'provides.'
  2. [§6, Interaction data] The sentence 'viewed weekly and daily explanations 13.042 and 10.33 times' appears to contain a typo ('13.042' should likely be '13.04'), and the means would be more interpretable with standard deviations or ranges.
  3. [§5.1 and §6] The power analysis targets a large effect (d=0.8) with a Bonferroni-corrected alpha of 0.05/3, but the study tests many outcome items across five hypothesis families. The rejected hypotheses (H1, H2a-H2c, H3e, H4b, H4c) should be described as 'not detected in this sample' rather than as evidence of no effect, given the limited power per item.
  4. [§5.3 and Table 2] The survey questions are described as 'based on the previous framework [75]' but the questionnaire itself is not included in full; consider adding it as an appendix or supplementary file so readers can assess item-construct alignment.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: empirical user-study evaluation; only a minor, non-load-bearing self-citation.

full rationale

PlanGlow is an empirical system evaluation rather than a derivation. The central claims are tested with user and expert ratings collected after the system was built; no model parameter was fitted to the outcome data, and no hypothesis was defined in terms of the observed ratings. The hypotheses (H1-H5) are operationalized through independent 7-point Likert and 5-point expert-rating items, and the paper honestly reports rejected hypotheses (H1, H2a-H2c/H2e-H2g, H3e, H4b-H4c, H5d). The mention of prior work [78] (Peerlens, which shares an author) as a basis for hypotheses is a self-citation, but it is not invoked as a uniqueness theorem, a fitted parameter, or a load-bearing justification for the empirical outcome; the comparison to GPT-4o and Khanmigo stands on its own measured data. The expert evaluation used a single rater for all plans, and the GPT-4o baseline is not feature-matched, but these are validity/confound concerns, not circularity. No equation, definition, or fitted parameter reduces to the paper's own inputs, so there is no specific circular step to exhibit. Score 1 reflects only the minor, non-load-bearing self-citation.

Assumptions & free parameters 2 free parameters · 8 assumptions · 0 invented entities

No mathematical derivation is involved. The claims rest on domain assumptions: survey items measure the intended constructs; expert rubric scores proxy educational quality; the baselines are fair comparators; n=24 with counterbalancing suffices; and ANOVA assumptions hold. Hand-chosen LLM sampling parameters are free design choices, not fitted to the evaluation data. No invented entities: PlanGlow is a software system, and explainability and controllability are borrowed constructs.

free parameters (2)
  • LLM sampling parameters for background descriptions = temperature 0.2, top_p 0.6, frequency_penalty 0.2, presence_penalty 0.1
    Hand-chosen values listed in Section 4 for generating background knowledge level descriptions; not fitted to the study data, but they are free design choices that affect generated text.
  • LLM sampling parameters for plan generation = temperature 0.0, top_p 0.8, frequency_penalty 0.2, presence_penalty 0.1
    Hand-tuned values for the three-step chain-of-thought plan generation described in Section 4; no ablation demonstrates that outputs depend on them.
assumptions (8)
  • domain assumption Learner background level can be mapped onto Benner's six-stage novice-to-expert framework and self-reported by users.
    Section 4 invokes framework [8] to define the background-level descriptions; personalization quality depends on this mapping being meaningful for the study population.
  • domain assumption Self-reported 7-point Likert ratings of usability, controllability, and explainability measure the constructs the hypotheses refer to.
    All H1-H4 conclusions rest on survey items in Table 2; no objective behavioral or learning-outcome measure is used for the main claims.
  • domain assumption Expert ratings of plan text using the Table 3 rubric proxy for educational quality of a study plan.
    Section 5.5: E1 scored simplified, text-only excerpts; no link between these ratings and actual learning outcomes is provided.
  • domain assumption The GPT-4o prompt-box baseline and Khanmigo's coach feature are fair comparators, so observed differences are due to PlanGlow's features rather than interface scaffolding.
    Section 5.2: the GPT-4o baseline has no structured UI, so it may underperform for reasons unrelated to explainability and controllability.
  • domain assumption Applying learning-science frameworks (Bloom, ZPD, andragogy, schema theory, metacognition) via prompts produces pedagogically better plans than generation without them.
    Section 4: the prompt design embeds these frameworks; expert-rated quality differences (H5) are attributed to this scaffolding, which is not ablated.
  • domain assumption Within-subject counterbalancing (6 orders x 4 participants) controls carryover and learning effects with n=24.
    Section 5.2; the sample is small and drawn from one university (Section 5.1), limiting statistical power for small effects.
  • standard math ANOVA with Bonferroni post-hoc tests is valid for the 7-point Likert data collected.
    Section 6 uses ANOVA on Likert items; no normality or variance-homogeneity checks are reported.
  • ad hoc to paper Assumed effect size d=0.8 in the power analysis justifies n=24.
    Section 5.1: the power analysis targets only large effects, so moderate or small true effects could be missed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PlanGlow: Personalized Study Planning with an Explainable and Controllable LLM-Driven System." pith.science (2026). https://pith.science/paper/2TX5HP4I

@misc{pith2026250412452,
  author       = {Pith},
  title        = {Pith review of: PlanGlow: Personalized Study Planning with an Explainable and Controllable LLM-Driven System},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2TX5HP4I}},
  note         = {Machine review of arXiv:2504.12452}
}
read the original abstract

Personal development through self-directed learning is essential in today's fast-changing world, but many learners struggle to manage it effectively. While AI tools like large language models (LLMs) have the potential for personalized learning planning, they face issues such as transparency and hallucinated information. To address this, we propose PlanGlow, an LLM-based system that generates personalized, well-structured study plans with clear explanations and controllability through user-centered interactions. Through mixed methods, we surveyed 28 participants and interviewed 10 before development, followed by a within-subject experiment with 24 participants to evaluate PlanGlow's performance, usability, controllability, and explainability against two baseline systems: a GPT-4o-based system and Khan Academy's Khanmigo. Results demonstrate that PlanGlow significantly improves usability, explainability, and controllability. Additionally, two educational experts assessed and confirmed the quality of the generated study plans. These findings highlight PlanGlow's potential to enhance personalized learning and address key challenges in self-directed learning.

Figures

Figures reproduced from arXiv: 2504.12452 by the authors.

Figure 1
Figure 1. PlanGlow contains two major components. (A) is the plan generation interface where users specify their learning preferences such as subject, prior knowledge, and available time. (A1) is the background level description. (B) is the generated plan interface that presents a personalized study plan, organized into weekly segments with a detailed daily breakdown. (B1) allows in-line editing of learning goals, background … view at source ↗
Figure 2
Figure 2. Detailed study plan of PlanGlow organizes each week into five days. (C1) explains the reasons for studying each day’s topic. (C2) lists learning objectives. (C3) allows users to explore additional resources via the button connecting to (C5). (C4) displays video resources with their status. A green check icon marks ‘Valid Resource’, while a red icon indicates ‘Invalid Resource’. (C5) displays 10 additional resources … view at source ↗
Figure 3
Figure 3. The workflow of PlanGlow progresses from left to right. The system begins by collecting user inputs through an input form, in-line editing, or chat interface to create an initial plan and describe the background knowledge level. The study plan is generated through three sequential steps using the OpenAI API: Initial Generation, Critique, and Improvement. The final plan incorporates comprehensive elements, including … view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: The interface of two systems are compared with [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Means and standard errors of GPT-4o-based system, [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Means and standard errors for each evaluation as [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

87 extracted references · 68 canonical work pages

  1. [1]

    Hasan Abu-Rasheed, Mohamad Hussam Abdulsalam, Christian Weber, and Mad- jid Fathi. 2024. Supporting student decisions on learning recommendations: An llm-based chatbot with knowledge graph contextualization for conversational explainability and mentoring. arXiv preprint arXiv:2401.08517 (2024)

  2. [2]

    Mohammad Amir Khusru Akhtar, Mohit Kumar, and Anand Nayyar. 2024. So- cially Responsible Applications of Explainable AI. In Towards Ethical and Socially Responsible Explainable AI: Challenges and Opportunities . Springer, 261–350

  3. [3]

    Farhan Ali, Doris Choy, Shanti Divaharan, Hui Yong Tay, and Wenli Chen. 2023. Supporting self-directed learning and self-assessment using TeacherGAIA, a generative AI chatbot application: Learning approaches and prompt engineering. Learning: Research and Practice 9, 2 (2023), 135–147

  4. [4]

    Hussam Alkaissi and Samy I McFarlane. 2023. Artificial hallucinations in Chat- GPT: implications in scientific writing. Cureus 15, 2 (2023)

  5. [5]

    Richard C Anderson. 2018. Role of the reader’s schema in comprehension, learning, and memory. In Theoretical models and processes of literacy . Routledge, 136–145

  6. [6]

    Nursultan Askarbekuly and Nenad Aničić. 2024. LLM examiner: automating assessment in informal self-directed e-learning using ChatGPT. Knowledge and Information Systems (2024), 1–18

  7. [7]

    Sinem Aslan, Lenitra M Durham, Nese Alyuz, Eda Okur, Sangita Sharma, Celal Savur, and Lama Nachman. 2024. Immersive multi-modal pedagogical conversa- tional artificial intelligence for early childhood education: An exploratory case study in the wild.Computers and Education: Artificial Intelligence 6 (2024), 100220

  8. [8]

    Patricia Benner. 1982. From novice to expert. AJN The American Journal of Nursing 82, 3 (1982), 402–407

Show all 87 references
  1. [9]

    Chantelle Bosch and Donnavan Kruger. 2024. AI chatbots as Open Educational Resources: Enhancing student agency and Self-Directed Learning. Italian Journal of Educational Technology (2024)

  2. [10]

    Svetlin Bostandjiev, John O’Donovan, and Tobias Höllerer. 2012. TasteWeights: a visual interactive hybrid recommender system. In Proceedings of the sixth ACM conference on Recommender systems . 35–42

  3. [11]

    Stefanie L Boyer, Diane R Edmondson, Andrew B Artis, and David Fleming. 2014. Self-directed learning: A tool for lifelong learning.Journal of marketing education PlanGlow: Personalized Study Planning with an Explainable and Controllable LLM-Driven System L@S ’25, July 21–23, 2...

  4. [12]

    AL Brown. 1978. Knowing when, where and how to remember. A problem of metacognition. In. R. Glaser (Ed.), Advances in instructional psychology (Vol. I). Hillsdale NJ Erlbaum 77 (1978), 165

  5. [13]

    Sheryl Burgstahler. 2009. Universal Design of Instruction (UDI): Definition, Principles, Guidelines, and Examples. Do-It (2009)

  6. [14]

    William Cain. 2024. Prompting change: exploring prompt engineering in large language model AI and its potential to transform education. TechTrends 68, 1 (2024), 47–57

  7. [15]

    Seth Chaiklin et al . 2003. The zone of proximal development in Vygotsky’s analysis of learning and instruction. Vygotsky’s educational theory in cultural context 1, 2 (2003), 39–64

  8. [16]

    Kathy Charmaz. 2006. Constructing grounded theory: A practical guide through qualitative analysis. sage

  9. [17]

    Blerta Abazi Chaushi, Besnik Selimi, Agron Chaushi, and Marika Apostolova

  10. [18]

    Li Chen and Feng Wang. 2017. Explaining recommendations based on feature sentiments in product reviews. In Proceedings of the 22nd international conference on intelligent user interfaces . 17–28

  11. [19]

    Constanţa Aurelia Chiţiba. 2012. Lifelong learning challenges and opportunities for traditional universities. Procedia-social and behavioral sciences 46 (2012), 1943–1947

  12. [20]

    Cristina Conati, Oswald Barral, Vanessa Putnam, and Lea Rieger. 2021. Toward personalized XAI: A case study in intelligent tutoring systems. Artificial intelli- gence 298 (2021), 103503

  13. [21]

    Coursera, Inc. 2024. Coursera. Online Learning Platform. https://www.coursera. org/ Accessed: 2024-12-10

  14. [22]

    Cecilia di Sciascio, Peter Brusilovsky, and Eduardo Veas. 2018. A study on user- controllable social exploratory search. In Proceedings of the 23rd International Conference on Intelligent User Interfaces . 353–364

  15. [23]

    Fengning Du. 2012. Using Study Plans to Develop Self-Directed Learning Skills: Implications from a Pilot Project. College Student Journal 46, 1 (2012)

  16. [24]

    Philippe Duchastel. 1983. Toward the ideal study guide: An exploration of the functions and components of study guides. British Journal of Educational Technology 14, 3 (1983), 216–231

  17. [25]

    Orpha K. Duell. 1986. MC Skills. In Cognitive Classroom Learning, Gary D. Phye and Thomas Andre (Eds.). Academic Press, Orlando, Florida

  18. [26]

    Elif Esiyok, Sahin Gokcearslan, and Kemal Gurkan Kucukergin. 2024. Acceptance of Educational Use of AI Chatbots in the Context of Self-Directed Learning with Technology and ICT Self-Efficacy of Undergraduate Students. International Journal of Human–Computer Interaction (2024), 1–10

  19. [27]

    Haoxiang Fan, Guanzheng Chen, Xingbo Wang, and Zhenhui Peng. 2024. Lesson- Planner: Assisting Novice Teachers to Prepare Pedagogy-Driven Lesson Plans with Large Language Models. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–20

  20. [28]

    Krzysztof Fiok, Farzad V Farahani, Waldemar Karwowski, and Tareq Ahram

  21. [29]

    Mehmet Firat. 2023. How chat GPT can transform autodidactic experiences and open education? (2023)

  22. [30]

    JH Flavell. 1976. Metacognitive aspects of problem solving. The nature of intelli- gence/Erlbaum (1976)

  23. [31]

    Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2024. Bias and fairness in large language models: A survey.Computational Linguistics (2024), 1–79

  24. [32]

    Xiu Guan, Xiang Feng, and AYM Islam. 2023. The dilemma and countermeasures of educational data ethics in the age of intelligence.Humanities and Social Sciences Communications 10, 1 (2023), 1–14

  25. [33]

    F Maxwell Harper, Funing Xu, Harmanpreet Kaur, Kyle Condiff, Shuo Chang, and Loren Terveen. 2015. Putting users in control of their recommendations. In Proceedings of the 9th ACM Conference on Recommender Systems . 3–10

  26. [34]

    Rex Hartson and Pardha S Pyla. 2012. The UX Book: Process and guidelines for ensuring a quality user experience . Elsevier

  27. [35]

    Jianxing He, Sally L Baxter, Jie Xu, Jiming Xu, Xingtao Zhou, and Kang Zhang

  28. [36]

    Wayne Holmes, Kaska Porayska-Pomsta, Ken Holstein, Emma Sutherland, Toby Baker, Simon Buckingham Shum, Olga C Santos, Mercedes T Rodrigo, Mutlu Cukurova, Ig Ibert Bittencourt, et al. 2022. Ethics of AI in education: Towards a community-wide framework. International Journal of ...

  29. [37]

    Dietmar Jannach, Sidra Naveed, and Michael Jugovac. 2017. User control in recommender systems: Overview and interaction challenges. In E-Commerce and Web Technologies: 17th International Conference, EC-Web 2016, Porto, Portugal, September 5-8, 2016, Revised Selected Papers 17 ...

  30. [38]

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. Comput. Surveys 55, 12 (2023), 1–38

  31. [39]

    Erik Jones and Jacob Steinhardt. 2022. Capturing failures of large language models via human cognitive biases. Advances in Neural Information Processing Systems 35 (2022), 11785–11799

  32. [40]

    Samad Kardan and Cristina Conati. 2015. Providing adaptive support in an interactive simulation for learning: An experimental evaluation. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems . 3671–3680

  33. [41]

    Judy Kay. 2001. Learner control. User modeling and user-adapted interaction 11 (2001), 111–127

  34. [42]

    Panayiota Kendeou and Victoria Johnson. 2023. The nature of misinformation in education. Current Opinion in Psychology (2023), 101734

  35. [43]

    Hassan Khosravi, Simon Buckingham Shum, Guanliang Chen, Cristina Conati, Yi-Shan Tsai, Judy Kay, Simon Knight, Roberto Martinez-Maldonado, Shazia Sadiq, and Dragan Gašević. 2022. Explainable artificial intelligence in education. Computers and Education: Artificial Intelligence...

  36. [44]

    Peter Kieseberg, Edgar Weippl, A Min Tjoa, Federico Cabitza, Andrea Campagner, and Andreas Holzinger. 2023. Controllable AI-An Alternative to Trustworthiness in Complex AI Systems?. In International Cross-Domain Conference for Machine Learning and Knowledge Extraction . Springer, 1–12

  37. [45]

    Rosemary Kim, Lorne Olfman, Terry Ryan, and Evren Eryilmaz. 2014. Leveraging a personalized system to improve self-directed learning in online educational environments. Computers & Education 70 (2014), 150–160

  38. [46]

    Simon Knight. 2020. Augmenting assessment with learning analytics. Re- imagining university assessment in a digital world (2020), 129–145

  39. [47]

    Simon Knight, Sophie Abel, Antonette Shibani, Yoong Kuan Goh, Rianne Conijn, Andrew Gibson, Sowmya Vajjala, Elena Cotos, Ágnes Sándor, and Simon Buck- ingham Shum. 2020. Are you being rhetorical? a description of rhetorical move annotation tools and open corpus of sample machi...

  40. [48]

    Malcolm Knowles. 1977. Adult learning processes: Pedagogy and andragogy. Religious Education 72, 2 (1977), 202–211

  41. [49]

    Malcolm S Knowles. 1978. Andragogy: Adult learning theory in perspective. Community College Review 5, 3 (1978), 9–20

  42. [50]

    2014.The adult learner: The definitive classic in adult education and human resource development

    Malcolm S Knowles, Elwood F Holton III, and Richard A Swanson. 2014.The adult learner: The definitive classic in adult education and human resource development . Routledge

  43. [51]

    Soila Lemmetty and Kaija Collin. 2020. Self-directed learning as a practice of workplace learning: Interpretative repertoires of self-directed learning in ICT work. Vocations and Learning 13, 1 (2020), 47–70

  44. [52]

    Belle Li, Curtis J Bonk, Chaoran Wang, and Xiaojing Kou. 2024. Reconceptualizing self-directed learning in the era of generative AI: An exploratory analysis of language learning. IEEE Transactions on Learning Technologies (2024)

  45. [53]

    Huiyong Li, Rwitajit Majumdar, Mei-Rong Alice Chen, Yuanyuan Yang, and Hiroaki Ogata. 2023. Analysis of self-directed learning ability, reading outcomes, and personalized planning behavior for self-directed extensive reading.Interactive Learning Environments 31, 6 (2023), 3613–3632

  46. [54]

    Su-Ting T Li, Daniel J Tancredi, John Patrick T Co, and Daniel C West. 2010. Factors associated with successful self-directed learning using individualized learning plans during pediatric residency. Academic pediatrics 10, 2 (2010), 124– 130

  47. [55]

    Fernando Antonio Flores Limo, David Raul Hurtado Tiza, Maribel Mamani Roque, Edward Espinoza Herrera, José Patricio Muñoz Murillo, Jorge Jinchuña Huallpa, Victor Andre Ariza Flores, Alejandro Guadalupe Rincón Castillo, Percy Fritz Puga Peña, Christian Paolo Martel Carranza, et...

  48. [56]

    Xi Lin. 2024. Exploring the role of ChatGPT as a facilitator for motivating self- directed learning among adult learners. Adult Learning 35, 3 (2024), 156–166

  49. [57]

    Huijin Lu, Maria Limniou, and Xiaojun Zhang. 2024. Exploring the metacognition of self-directed informal learning on social media platforms: taking time and social interactions into consideration. Education and Information Technologies (2024), 1–28

  50. [58]

    Rose Luckin and Wayne Holmes. 2016. Intelligence unleashed: An argument for AI in education. (2016)

  51. [59]

    Aniek F Markus, Jan A Kors, and Peter R Rijnbeek. 2021. The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies. Journal of biomedical informatics 113 (2021...

  52. [60]

    Massachusetts Institute of Technology. 2024. MIT OpenCourseWare. Open Educational Resource. https://ocw.mit.edu/ Accessed: 2024-12-10

  53. [61]

    Bahar Memarian and Tenzin Doleck. 2023. ChatGPT in education: Methods, potentials and limitations. Computers in Human Behavior: Artificial Humans (2023), 100022

  54. [62]

    Matthew B Miles, A Michael Huberman, and Johnny Saldaña. 2014. Qualitative data analysis: A methods sourcebook. 3rd. L@S ’25, July 21–23, 2025, Palermo, Italy Chun et al

  55. [63]

    Brent Mittelstadt. 2019. Principles alone cannot guarantee ethical AI. Nature machine intelligence 1, 11 (2019), 501–507

  56. [64]

    Vahid Mohammadi, Amir Masoud Rahmani, Abdullah Mohammad Darwesh, et al

  57. [65]

    University of Arkansas. n.d.. Using Bloom’s Taxonomy. https://tips.uark.edu/ using-blooms-taxonomy/ Accessed: 2024-06-19

  58. [66]

    Linnea Öhlund. 2020. Valuable Visuals: Defining a design space for presenting medical results

  59. [67]

    Jeroen Ooge, Shotallo Kato, and Katrien Verbert. 2022. Explaining recommen- dations in e-learning: Effects on adolescents’ trust. In Proceedings of the 27th International Conference on Intelligent User Interfaces . 93–105

  60. [68]

    Human-centric Computing and Information Sciences 9 (2019),

    Trust-based recommendation systems in Internet of Things: a systematic literature review. Human-centric Computing and Information Sciences 9 (2019),

  61. [69]

    https://doi.org/10.1186/s13673-019-0183-8

  62. [70]

    Ryan Rafiola, Punaji Setyosari, Carolina Radjah, and M Ramli. 2020. The effect of learning motivation, self-efficacy, and blended learning on students’ achievement in the industrial revolution 4.0. International Journal of Emerging Technologies in Learning (iJET) 15, 8 (2020), 71–82

  63. [71]

    Arvind Satyanarayan and Graham M. Jones. 2024. Intelligence as Agency: Evaluating the Capacity of Generative AI to Empower or Constrain Human Action. An MIT Exploration of Generative AI (mar 27 2024). https://mit- genai.pubpub.org/pub/94y6e0f8

  64. [72]

    Jordan Richard Schoenherr, Roba Abbas, Katina Michael, Pablo Rivas, and Theresa Dirndorfer Anderson. 2023. Designing AI using a human-centered approach: Explainability and accuracy toward trustworthiness. IEEE Transactions on Technology and Society 4, 1 (2023), 9–23

  65. [73]

    Hyanghee Park and Daehwan Ahn. 2024. The Promise and Peril of ChatGPT in Higher Education: Opportunities, Challenges, and Design Implications. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–21

  66. [74]

    Kaśka Porayska-Pomsta, Wayne Holmes, and Selena Nemorin. 2023. The ethics of AI in education. In Handbook of Artificial Intelligence in Education . Edward Elgar Publishing, 571–604

  67. [75]

    Zhijie Wang, Yuheng Huang, Da Song, Lei Ma, and Tianyi Zhang. 2024. PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–21

  68. [76]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837

  69. [77]

    Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–22

  70. [78]

    Donggil Song and Curtis J Bonk. 2016. Motivational factors in self-directed informal learning from online learning resources. Cogent Education 3, 1 (2016), 1205838

  71. [79]

    Eric J Topol. 2019. High-performance medicine: the convergence of human and artificial intelligence. Nature medicine 25, 1 (2019), 44–56

  72. [80]

    Ling Zhang, James D Basham, and Sohyun Yang. 2020. Understanding the implementation of personalized learning: A research synthesis. Educational research review 31 (2020), 100339

  73. [81]

    Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A Smith. 2023. How language model hallucinations can snowball. arXiv preprint arXiv:2305.13534 (2023)

  74. [82]

    Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for large lan- guage models: A survey. ACM Transactions on Intelligent Systems and Technology 15, 2 (2024), 1–38

  75. [83]

    Meng Xia, Mingfei Sun, Huan Wei, Qing Chen, Yong Wang, Lei Shi, Huamin Qu, and Xiaojuan Ma. 2019. Peerlens: Peer-inspired interactive learning path planning in online question pool. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–12

  76. [84]

    Ling Xu. 2020. The dilemma and countermeasures of AI in educational application. In Proceedings of the 2020 4th international conference on computer science and artificial intelligence. 289–294

  77. [2019]

    Nature medicine 25, 1 (2019), 30–36

    The practical implementation of artificial intelligence technologies in medicine. Nature medicine 25, 1 (2019), 30–36

  78. [2022]

    The Journal of Defense Modeling and Simulation 19, 2 (2022), 133–144

    Explainable artificial intelligence for education and training. The Journal of Defense Modeling and Simulation 19, 2 (2022), 133–144

  79. [2023]

    In World Conference on Explainable Artificial Intelligence

    Explainable artificial intelligence in education: A comprehensive review. In World Conference on Explainable Artificial Intelligence. Springer, 48–71

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.