{"id":"e15b9ba3-2e20-44d8-acc2-29ec117b4731","arxiv_id":"2504.12452","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"An LLM-driven study planning system that adds hierarchical explanations and real-time user controls outperformed ChatGPT and Khanmigo on user-rated explainability, controllability, and expert-assessed plan quality.","lead":"PlanGlow is a new AI-based study planner that adds explanations for every recommended topic and lets learners swap resources and edit plans in place. In a 24-person study, students preferred it over ChatGPT and Khanmigo, rating it higher for control, explanations, and plan quality.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Attribution of PlanGlow's user-perceived gains to explainability/controllability is confounded by the feature-unmatched GPT-4o prompt-box baseline; an ablation/feature-matched control is needed to support the headline claim.","rationale":"The paper is a competent systems contribution with honest null results. The load-bearing issue is the step from observed differences to the design-feature attribution. The primary comparisons are between a full structured web app and a bare prompt box/commercial chat, so the significant H3/H4 items largely measure the presence of artifacts (validation icons, explanation blocks, structured forms) rather than the quality or causal effect of the explanation/control design. That does not make the paper's system-level conclusion false; it makes the mechanistic conclusion underdetermined. The expert plan-quality result has a separate single-rater presentation confound, but it is secondary because the central user-facing claims rest on H3/H4. The reader already marked the paper CONDITIONAL for essentially this reason, so the verdict should remain unchanged. No ad hominem or internal-contradiction claim is intended; this is a standard construct-validity and baseline-matching concern.","tokens_in":17764,"tokens_out":6619,"duration_ms":73636,"concrete_test":"Run an ablation study with the same PlanGlow interface in three arms: full PlanGlow, PlanGlow-NoExplain (same layout, controls, and resources, but rationale/objective/connection text removed or replaced with neutral placeholders), and PlanGlow-NoControl (same explanations and layout, but in-line editing, video replacement, and validation icons disabled), counterbalanced with n=24 and the same survey. If full PlanGlow does not significantly outperform both ablated arms on H3b/H3d/H4d-H4g, the headline advantages are attributable to the UI scaffold rather than to the explainability/controllability features.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that explainability and controllability features drive PlanGlow's improvements is not supported by the current comparison. In §5.2, the GPT-4o baseline is a bare text box; Khanmigo is a commercial chat interface with different layout. Several significant results therefore measure feature presence rather than design quality. H3d asks about 'reliability of resource validation' and only PlanGlow displays validation icons (Figure 2 C4); H4d-H4g ask about rationales, goal alignment, task connections, and informed decisions, and GPT-4o's raw text output contains no explicit explanation blocks; H2d is 'functional integration,' which a prompt box cannot demonstrate. The paper's own DC2 reports that in-line editing and chat modifications were rarely used, so the controllability ratings reflect available affordances rather than exercised control. The expert-quality result has the same problem: §5.5 used one rater (E1) with no inter-rater reliability, and PlanGlow was evaluated from a partial excerpt while E1 was told the full plan and validated resources existed, introducing differential presentation and expectation bias. The abstract's phrase 'two educational experts assessed' overstates this protocol. The experiment supports 'a purpose-built plan UI is preferred to a prompt box,' but not the attribution to the named XAI/control constructs. This is a validity limitation, not an internal inconsistency; the honest null results and high explanation engagement are useful evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents PlanGlow, a web-based LLM-driven study planning system with structured input forms, generated plans with explanations (rationale, learning objectives, task connections), YouTube API-based resource validation, in-line editing, and a chat feature. The authors conducted a formative survey (n=28) and interviews (n=10 plus one educational researcher) to derive four design requirements, then ran a within-subject experiment with 24 students comparing PlanGlow to a GPT-4o prompt-box baseline and Khan Academy's Khanmigo. Outcome measures were 7-point Likert ratings on performance, usability, controllability, and explainability (H1-H4) and expert ratings of plan quality (H5). Results show no significant performance differences (H1 rejected), a single usability win on functional integration (H2d), significant controllability gains on plan-generation ease vs Khanmigo, efficiency, resource search vs Khanmigo, and resource validation reliability (H3a-H3d), significant explainability gains on conciseness vs Khanmigo and on rationale, goal alignment, task connections, and informed decisions (H4d-H4g), and expert-rated quality advantages on multiple H5 sub-hypotheses. The abstract claims that PlanGlow 'significantly improves usability, explainability, and controllability,' and 83.3% of participants ranked PlanGlow as their top choice.","tokens_in":17982,"tokens_out":3734,"duration_ms":37485,"significance":"If the reported results are taken at face value, this is a useful empirical evaluation of an LLM-based tool for the planning stage of self-directed learning, an area the paper correctly identifies as underexplored. The manuscript has several strengths: a counterbalanced within-subject design, pre-specified hypotheses, Bonferroni post-hoc corrections, detailed interaction logs, honest reporting of rejected hypotheses (H1, most of H2, H3e, H4b, H4c, H5d), and publicly available code and evaluation materials. The system design, including chain-of-thought generation with a critique and improvement step and YouTube API-based resource validation, is concrete and reproducible. The main limitation is interpretive: the measured advantages are not cleanly attributable to the named explainability and controllability constructs because the comparators differ in many interface dimensions. The honest null results and high engagement with explanation features are nevertheless valuable for future XAI-in-education work.","major_comments":[{"comment":"The attribution of PlanGlow's advantages to explainability and controllability is confounded by the feature-unmatched baseline. The GPT-4o comparator is a bare text box with no structured form, plan visualization, editing affordances, resource display, or explicit explanation blocks. Because H3d (resource validation reliability), H4d-H4g (rationale, goal alignment, task connections, informed decisions), and H2d (functional integration) are measured at the feature level, the significant differences could simply reflect the presence of any purpose-built interface rather than the quality of the explainability and control designs. To support the headline claim, the paper would need a feature-matched control or ablation (e.g., PlanGlow with explanations and controls disabled). As it stands, the evidence supports a system-level comparison ('a purpose-built plan UI was preferred to a prompt box'), not a causal attribution to the named constructs.","section":"§5.2, Figure 4; §6, H2-H4 results"},{"comment":"The expert evaluation protocol is not as described in the abstract. Section 5.5 states that E1 evaluated all generated plans, while E2 only collaborated on developing the evaluation criteria; the abstract and Introduction's claim that 'two educational experts assessed' the plans overstates the protocol. There is also no inter-rater reliability coefficient. In addition, PlanGlow plans were presented as simplified Week-1/Day-1 excerpts while E1 was told that full explanations existed and resources had been validated, whereas the GPT-4o and Khanmigo plans were presumably presented in full text. This differential presentation and expectation cue could plausibly bias the H5a-H5c and H5e ratings. Please report the exact presentation format for all three systems, justify the belief that the excerpts are representative, and either provide a second independent rater or clearly frame H5 as a single-rater exploratory assessment.","section":"§5.5 and Table 3; abstract; §1"},{"comment":"The abstract's claim that PlanGlow 'significantly improves usability, explainability, and controllability' is not supported by the reported statistics. Usability was significant only for H2d (functional integration); H2a-H2c and H2e-H2g were rejected. Explainability was significant for conciseness only against Khanmigo and was not significant for accuracy (H4b) or relevance (H4c). The Conclusion's statement that PlanGlow 'significantly outperformed two baselines' is likewise overly broad given the many non-significant comparisons. The abstract and conclusion should be revised to state precisely which sub-hypotheses were supported, or the paper should present a composite measure (if justified) before claiming overall improvements in those dimensions.","section":"Abstract; §8 Conclusion"}],"minor_comments":[{"comment":"The text contains garbled notation and a duplicated phrase: 'the \"bulb\" ♂lightbulbicon' should be cleaned up, and 'the system provides a clear visual indicator provides real-time feedback' repeats the verb 'provides.'","section":"§4, Figure 1 caption and text"},{"comment":"The sentence 'viewed weekly and daily explanations 13.042 and 10.33 times' appears to contain a typo ('13.042' should likely be '13.04'), and the means would be more interpretable with standard deviations or ranges.","section":"§6, Interaction data"},{"comment":"The power analysis targets a large effect (d=0.8) with a Bonferroni-corrected alpha of 0.05/3, but the study tests many outcome items across five hypothesis families. The rejected hypotheses (H1, H2a-H2c, H3e, H4b, H4c) should be described as 'not detected in this sample' rather than as evidence of no effect, given the limited power per item.","section":"§5.1 and §6"},{"comment":"The survey questions are described as 'based on the previous framework [75]' but the questionnaire itself is not included in full; consider adding it as an appendix or supplementary file so readers can assess item-construct alignment.","section":"§5.3 and Table 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for L@S and contains a carefully reported empirical study with useful honest null results. The main problems are the mismatch between the abstract's broad claims and the evidence, and the confounded baseline design. These are fixable by reframing the contribution as a system-level comparison and adding an ablation or feature-matched condition, or by explicitly limiting the causal claims. I do not see grounds for rejection, but the revision must address the attribution issue and the overstated expert-evaluation claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a competent HCI systems paper about the planning stage of self-directed learning, a genuinely under-served part of the pipeline. The formative work is real—28 surveys, 10 interviews, one education expert—and it produces four concrete design requirements. PlanGlow implements those requirements with hierarchical explanations, YouTube resource validation, in-line editing, and a three-step chain-of-thought generation process. The evaluation is a counterbalanced within-subject design with 24 participants, a power analysis, Bonferroni-corrected ANOVAs, and—to the authors' credit—a lot of honest null results: performance differences were not significant, most usability hypotheses were rejected, and explainability accuracy/relevance did not separate from baselines. They also report interaction logs showing heavy use of explanations and light use of the editing and chat features. That is useful, reproducible evidence.\n\nThe soft spots are real but specific. The abstract says PlanGlow 'significantly improves usability, explainability, and controllability,' which the data do not support as a blanket statement. What actually held up was functional integration, several controllability items (especially resource validation and search), and several explainability items about rationale, goal alignment, and task connections. The bigger issue is attribution. The GPT-4o baseline is a bare text box, so many significant items—functional integration, resource validation, rationale explanations—could simply reflect having a purpose-built UI rather than the specific explainability/controllability constructs the paper claims. Khanmigo is also a different interface, not a feature-matched control. The expert scoring has a similar problem: one rater, no inter-rater reliability, and PlanGlow plans were shown as partial excerpts while the rater was told the full plan and validated resources existed. That can bias the H5 results in PlanGlow's favor. The paper's own limitation section acknowledges resource breadth and YouTube-only design, but it does not confront the baseline confound head-on.\n\nWho should read this: anyone building LLM-based study-planning tools or studying explainability/controllability in educational AI will get concrete design ideas and a template for an honest within-subject evaluation. The citation pattern looks fine, and the code and OSF links are a plus. It deserves a serious referee, but the referee should push for either a feature-matched control or a sharpened claim about what was actually compared. I would not accept the headline as written; I would accept the underlying system and evaluation as a solid contribution after revision.","headline":"Solid, honest systems paper whose headline claims run ahead of a feature-unmatched baseline and single-rater expert scoring; the design patterns are still worth engaging with.","tokens_in":18575,"tokens_out":1247,"would_cite":true,"duration_ms":16193,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PlanGlow claims that adding explanations and user control to an LLM study planner beats a bare GPT-4o prompt box and Khanmigo on usability, explainability, controllability, and expert-rated plan quality.","keywords":["Self-directed learning","Personalized learning","Explainable AI","Controllable AI","Large language models","Study planning","User study","AI hallucination"],"falsifier":"Run the same 24-participant within-subject comparison against a feature-matched control: the same structured form, editing, and resource list as PlanGlow, but with the rationale boxes, explanation toggles, and video-validation badges removed. If users no longer rate the control lower on H3b, H3d, and H4d–H4g, then the features are not the driver; if they do, the paper's interpretation holds. For the expert ratings, two independent raters scoring full plans in identical plain-text formatting would test whether the H5 advantage survives presentation bias.","tokens_in":17476,"feed_emoji":"📚","tokens_out":13107,"duration_ms":123781,"temperature":0.7,"pith_summary":"PlanGlow is an LLM-based study-planning system built to test a specific claim: that learners' problems with AI-generated study plans—opaque reasoning, hallucinated or mismatched resources, and difficulty adjusting the plan—can be addressed by design features rather than by a better model. The paper reports a within-subject study in which 24 learners compared PlanGlow with a bare GPT-4o prompt box and with Khanmigo, and one educator scored the resulting plans against criteria developed with a second educator. PlanGlow scored significantly higher on functional integration, efficient plan generation, reliable resource validation, and on explanations of recommendation rationale, goal alignment, and task connections, and it received significantly higher expert ratings on objectives, timelines, resources, and pedagogical soundness. Raw performance and most usability items showed no significant difference. If the claim is right, the takeaway for AI learning tools is that explanations and user control, not the underlying model, drive measurable gains in perceived plan quality at the planning stage of self-directed learning.","feed_headline":"Explanations and control beat bare-bones GPT-4o for study plans","feed_subtitle":"In a 24-person test, the explainable planner beat a text-box AI and Khanmigo on rationale, goals, and resources.","key_machinery":"The load-bearing object is PlanGlow's plan-generation and presentation pipeline: a structured input form collects subject, goals, background level, duration, and daily availability; a three-step chain-of-thought procedure (initial generation, critique, improvement) produces the plan; a layered interface exposes weekly overviews, daily breakdowns, rationale boxes, and video-validation status; and in-line editing, chat, and resource replacement give users control. The features that carry the argument are the explanation panels (rationale, objectives, connections) and the validated resource list, because they distinguish PlanGlow from chat-based baselines and map directly to the hypotheses that showed significant gains.","core_discovery":"On the paper's own terms, the discovery is that wrapping an LLM study-plan generator in explanation and control features moves user-perceived and expert-rated plan quality. In a within-subject comparison with a GPT-4o prompt box and Khanmigo, PlanGlow scored significantly higher on functional integration (H2d), efficient generation of the desired plan (H3b), and reliable resource validation (H3d), and, against Khanmigo, on easy plan generation (H3a) and straightforward alternative-resource search (H3c). It also scored significantly higher on explaining the rationale behind recommendations, aligning plans with goals, clarifying daily-weekly task connections, and enabling informed decisions (H4d–H4g), and on concise explanations versus Khanmigo (H4a). Expert ratings favored PlanGlow on learning objectives, timelines, resources, and pedagogical soundness against both baselines (H5a–H5c, H5e), while progress monitoring was not significantly better than GPT-4o (H5d). Overall performance, most usability items, and explanation accuracy or relevance did not differ significantly, and 83.3% of participants ranked PlanGlow as their top choice.","pith_inferences":["The paper leaves open whether users rate explanations highly because they genuinely inform decisions or because they signal care and structure; a follow-up that measures decision quality, not just ratings, would separate the two.","Because the GPT-4o comparator was a bare text box, part of the gap may come from having any purpose-built interface; a feature-matched control that removes only the rationale panels would isolate the explainability effect.","The low observed use of in-line editing and chat suggests that the visible presence of control may matter more than actual manipulation in a single session, so whether control features pay off over weeks of self-directed learning is a natural longitudinal test.","The absence of significant gains on explanation accuracy or relevance and on progress monitoring implies PlanGlow shifts perceived quality and planning structure without yet demonstrating improved factual reliability or execution support; coupling the system with retrieval or assessment tools would test whether those gaps close."],"forward_implications":["If the results hold, adding structured rationale and resource validation to an LLM planner is enough to move user-perceived plan quality even when the underlying model stays the same.","Learning platforms that already offer LLM chat could adopt the form-plus-editing-plus-explanation pattern without changing their model.","Because the usability win was functional integration rather than general ease of use, the design lesson is feature coherence, not overall polish.","Planning-stage tools for self-directed learners can expect expert-rated objectives, timelines, resources, and pedagogy to improve when plans carry explicit rationale, and should not rely on raw model output alone."],"supporting_citations":[{"why":"Supplies the GPT-4o lesson-planning system used as the primary comparison baseline.","marker":"[27]"},{"why":"Chain-of-thought prompting is the basis for PlanGlow's initial-critique-improvement generation pipeline.","marker":"[76]"},{"why":"The user-experience survey items on controllability and explainability are adapted from this prior work.","marker":"[75]"},{"why":"The performance and usability hypotheses are framed against this earlier learning-path planning work, shaping the survey items.","marker":"[78]"},{"why":"The learning-objectives taxonomy used to structure plan objectives that the expert rater scored in H5a.","marker":"[65]"},{"why":"Supplies the evaluation criteria the educator used to score plan objectives, timelines, resources, and pedagogy (H5).","marker":"[23, 54]"}],"fun_headline_variants":["Explainable study planner tops GPT-4o and Khanmigo","Users prefer explainable study plans over bare GPT-4o","Explainable and controllable planner beats Khanmigo","LLM study planner with explanations wins user top choice","Explainable study plans beat plain GPT-4o and Khanmigo"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on the assumption that PlanGlow's measured advantages come from its explanation and control features rather than from simply having a purpose-built interface, since its main comparator was a bare text box and only one expert rated simplified plan excerpts.","fun_headline_variants_meta":{"raw":{"variants":["Explainable study planner tops GPT-4o and Khanmigo","Users prefer explainable study plans over bare GPT-4o","Explainable and controllable planner beats Khanmigo","LLM study planner with explanations wins user top choice","Explainable study plans beat plain GPT-4o and Khanmigo"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000662,"raw_usage":{"total_tokens":3042,"prompt_tokens":981,"completion_tokens":2061,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":1977}},"tokens_in":597,"tokens_out":2061,"duration_ms":15609,"temperature":1.0,"reasoning_tokens":1977,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:31:24.147848+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 24-participant within-subject comparison against a feature-matched control: the same structured form, editing, and resource list as PlanGlow, but with the rationale boxes, explanation toggles, and video-validation badges removed. If users no longer rate the control lower on H3b, H3d, and H4d–H4g, then the features are not the driver; if they do, the paper's interpretation holds. For the expert ratings, two independent raters scoring full plans in identical plain-text formatting would test whether the H5 advantage survives presentation bias.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the GPT-4o lesson-planning system used as the primary comparison baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The performance and usability hypotheses are framed against this earlier learning-path planning work, shaping the survey items."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The learning-objectives taxonomy used to structure plan objectives that the expert rater scored in H5a."}],"review_version":1}