{"id":"b934402b-2232-4b18-b52c-d48902d04ff5","arxiv_id":"2607.03303","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Socratic LLM tutoring, not prompt-refinement coaching, later yields higher learning gains and more understanding-driven prompting with unconstrained LLMs in a robotics course.","lead":"A classroom experiment found that a Socratic LLM tutor produced higher later learning gains and more understanding-oriented prompting than a prompt-refinement tutor, even though both looked similar during guided use. The result matters because it suggests how we scaffold AI tutors may shape durable student habits with unconstrained LLMs.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The S3 learning-gain interaction rests on a pre-test imbalance that the mixed model does not fully neutralize, and the transfer claim is further weakened by group-level partial observability.","rationale":"The Reader correctly flags the causal leap from SG assignment to later understanding-driven prompting under partial observability and self-selection. That concern is load-bearing, but the S3 learning-gain interaction itself is equally fragile because of the documented pre-test imbalance and modest model fit; both must hold for the developmental story. The concrete checks above directly test whether those two results survive baseline equalization and complete-trace restriction. If they do, the CONDITIONAL verdict remains appropriate and confidence can rise; if either fails, the claim should be treated as exploratory only. No stronger objection (e.g., internal inconsistency of the coding scheme or tutor implementation) is warranted from the text. Thus the Reader’s verdict is left unchanged while the weakest link is sharpened.","tokens_in":12649,"tokens_out":613,"duration_ms":5989,"concrete_test":"Re-fit the S3 repeated-measures model after (a) residualizing post-test on pre-test within condition or (b) exact matching / propensity weighting on pre-test; if the Time×Condition interaction falls below p=0.05 or |β|<0.4, the within-intervention learning claim is unreliable. Separately, re-estimate the project understanding regression using only groups with complete intervention-student chatbot traces; if the Understanding-driven coefficient loses significance, the transfer claim is confounded.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central developmental claim (SG yields higher later learning gains and more understanding-driven unconstrained prompting) depends on two linked results. First, the S3 pre–post interaction (Table 1: condition main effect β=−0.52, p=0.02; interaction β=0.72, p=0.01) is the only within-intervention learning difference. Yet Section 3.4 already notes a marginal pre-test imbalance favoring PR (U=380, p=0.06), and the model reports only a modest conditional R² (0.23). With n≈65 and no reported sensitivity check that equalizes baselines or drops the lowest-pre SG students, the interaction can be produced by regression-to-the-mean or differential ceiling rather than by the tutor. Second, the transfer result (χ²=8.14 on project clusters; group-level regression linking Understanding-driven counts to oral scores) is estimated on 30 groups whose composition mixes SG/PR/non-intervention students and whose prompting traces are only partially observed (Section 3.5). The paper’s own regression therefore cannot cleanly attribute later prompting style or understanding to prior SG assignment versus self-selection into chatbot use or unobserved group dynamics. If either the S3 interaction or the project association is artifactual, the developmental claim collapses.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper compares two LLM tutors in a graduate mobile-robotics course: a Socratic-Guidance (SG) tutor that answers with reflective questions and a Prompt-Refinement (PR) tutor that gates assistance until prompts meet clarity and learning-intent criteria. In a between-subjects intervention (n=66 across three labs), the tutors produce similar prompting clusters and comparable task performance by later sessions, but SG yields a significant pre–post learning interaction in session S3. In a subsequent unconstrained-LLM project phase (n=52), former SG students more often fall into an understanding-driven high-quality prompting cluster, and group-level counts of such prompters predict higher oral-exam understanding. The authors conclude that response-level Socratic scaffolding, despite lower perceived efficiency, better supports durable capacity to learn with unconstrained LLMs.","tokens_in":13031,"tokens_out":1166,"duration_ms":9555,"significance":"If the developmental transfer claim holds, the work supplies concrete design evidence that pedagogically shaping LLM responses (Socratic questioning) can outperform teaching prompt form alone for later independent LLM use—an important and under-tested distinction for AIED and computing-education design. Strengths include random assignment to tutors, pre/post tests and practice scores with substantial inter-rater reliability, multi-dimensional prompt coding with reported Fleiss’ κ, spectral clustering of prompting patterns, and mixed-effects modeling of learning. The two-phase design (scaffolded labs then unconstrained project) is a genuine contribution relative to single-session prompt-interface studies. The result is therefore of real interest to the field even if some causal links require tighter checks.","major_comments":[{"comment":"§4.2 / Table 1: The sole within-intervention learning difference is the S3 Time×Condition interaction (β=0.72, p=0.01), yet §3.4 already reports a marginal pre-test imbalance favoring PR (U=380, p=0.06) and the model shows only modest conditional R² (0.23). With n≈65 and no reported sensitivity analysis that equalizes baselines, drops low-pre SG students, or checks ceiling/regression-to-the-mean, the interaction is not yet load-bearing for the claim that SG produces superior later learning gains.","section":"§4.2 Table 1"},{"comment":"§3.5 and §4.4: The transfer claim rests on project-phase cluster differences (χ²=8.14) and a group-level regression of oral understanding on counts of Understanding-driven prompters. Scores are assigned to 30 mixed groups (SG/PR/non-intervention students), chatbot use is voluntary and only partially observed, and the regression cannot cleanly separate prior SG assignment from self-selection into chatbot use or unobserved group dynamics. This is the weakest link in the developmental claim and needs either individual-level outcomes, fuller trace coverage, or stronger identification arguments.","section":"§3.5 §4.4"},{"comment":"§4.1 vs §4.4: During guided use the tutors produce statistically indistinguishable prompting clusters, yet after scaffold removal SG students disproportionately adopt understanding-driven patterns. The paper treats this as evidence of a developmental influence, but offers no process account (e.g., changes in self-authored text proportion, dialogue length, or metacognitive language) that would make the delayed effect more than a post-hoc association. Without such bridging evidence the central design recommendation remains under-supported.","section":"§4.1 §4.4"}],"minor_comments":[{"comment":"§3.5: Free parameters of the clustering pipeline (Gaussian kernel σ=0.0001, Silhouette-selected feature aggregation of conceptual∪understanding and clarity∪granularity) should be justified or subjected to a brief robustness check; small cell sizes (e.g., n=8 Understanding-driven in S3) make cluster labels sensitive to these choices.","section":"§3.5"},{"comment":"§3.4: Post-test construction differs from pre-test (MCQs + reweighted sequencing vs open algorithmic explanation); a short note on score comparability or standardization would help readers interpret pre–post gains.","section":"§3.4"},{"comment":"§4.5: Perception results are reported with Kruskal–Wallis H and corrected p-values; effect-size interpretation is given, but the direction of the efficiency/learning trade-off would be clearer with medians or rank-biserial values by condition.","section":"§4.5"},{"comment":"Figure 2 and Figure 3 captions state that only significant interactions are indicated; adding exact n per condition/cluster in the panels would aid interpretation given the small cells.","section":"Figures 2–3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid empirical AIED study with real methodological care (IRR, mixed models, two-phase design). The developmental claim is interesting but currently rests on a fragile S3 interaction and a partially observed group-level transfer analysis; major revision that either strengthens identification or softens causal language would make it a clear contribution. Fit for a serious AIED / CS-education venue is good if the statistical concerns are addressed."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is the design, not a single p-value: they ran the same graduate robotics course with two LLM tutors—Socratic response scaffolding versus a prompt-refinement gate—then stripped the scaffolds and watched unconstrained use on a project. That two-phase contrast is the actual addition. Most of the Socratic-tutor literature stops at guided sessions; most of the prompt-engineering literature never checks transfer. This paper does both in one course.\n\nWhat they do well is the craft of the classroom study. Random assignment, pre/post with solid IRR, prompt coding with reported kappa, mixed-effects models, and an honest perception result: students liked the PR tutor more and still learned less by S3. The cluster finding that understanding-driven high-quality prompts track learning (not task score) lines up with their prior work and is cleanly reported. No circularity games; the constructs are measured, not redefined into existence.\n\nThe soft spots are real and about the size the stress-test says. The only within-intervention learning win is the S3 time×condition interaction, and S3 already had a marginal pre-test tilt against SG. With n≈65 and conditional R² around 0.23, regression-to-the-mean is a live alternative. The transfer claim is weaker still: project understanding is group-level, groups mix SG/PR/non-intervention students, and prompting traces are only partially observed. Their own regression cannot cleanly separate prior tutor from self-selection into chatbot use. Cluster cells are small. No public code or data. None of that kills the paper; it means the developmental claim is directional evidence, not settled design law.\n\nThis is for people building educational LLM tutors and for computing-education researchers who care about durable habits rather than session scores. It deserves a serious referee, not a desk reject—revise for baseline sensitivity, individual-level transfer measures, and clearer limits on the group model. I would bring it to an AI-ed reading group and cite the design and the transfer pattern with the caveats attached. Engage with it; do not over-claim from it.","headline":"Solid classroom head-to-head of Socratic vs prompt-refinement tutors; the delayed-learning and transfer story is the real contribution, but both rest on modest n and messy baselines.","tokens_in":13653,"tokens_out":535,"would_cite":true,"duration_ms":11474,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Socratic LLM tutors raise later learning gains and transfer understanding-driven prompting once scaffolds are removed, even though students rate them as less efficient.","keywords":["Large Language Models","Computing Education","Socratic tutoring","prompt refinement","scaffolding","prompting strategies","transfer","learning gains"],"falsifier":"A larger randomized follow-up that forces full individual-level tracing of unconstrained LLM use and individual (not group oral) understanding scores: if Socratic assignment no longer predicts understanding-driven clusters or understanding once self-selection and composition are blocked, the transfer claim fails.","tokens_in":13531,"feed_emoji":"🎓","tokens_out":574,"duration_ms":4971,"temperature":0.7,"pith_summary":"This paper asks whether students learn better long-term ways of using large language models when the tutor structures dialogue through Socratic questions or when it trains them to write better prompts. In a graduate robotics course, students used one of two scaffolded tutors across lab sessions, then switched to an unconstrained course chatbot for a team project. During guided use the two tutors produced similar prompt patterns and eventually similar task scores, but by the last lab the Socratic group showed larger learning gains. After scaffolds were removed, those same students were more likely to ask understanding-oriented questions, and groups with more such prompters scored higher on project understanding. The practical point is that response-side reflective scaffolding can shape durable LLM habits more than prompt-side refinement, even when learners feel the reflective tutor is less efficient.","feed_headline":"Socratic LLM tutors build better free-use prompting habits","feed_subtitle":"They lag in perceived efficiency but raise later learning and understanding-driven questions once scaffolds vanish.","key_machinery":"The two-phase contrast between a Socratic-Guidance tutor (response-side reflective questioning) and a Prompt-Refinement tutor (learner-side prompt quality gates), followed by unconstrained LLM use, with spectral clustering of prompts into understanding-driven versus implementation/debugging patterns that are then linked to pre-post learning and group oral understanding.","core_discovery":"Relative to a Prompt-Refinement tutor, a Socratic-Guidance tutor produces higher learning gains by later guided sessions and, once scaffolds are removed, higher rates of understanding-driven prompting that predict higher project understanding with an unconstrained LLM.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Socratic tutors train lasting free-use prompting habits","Questioning tutors beat prompt tips for later LLM skills","Socratic guidance raises later gains and smarter free LLM use","Dialogic tutors build understanding-driven prompting after scaffolds","Socratic LLM scaffolding yields better independent prompting"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That later differences in unconstrained prompting and group project understanding are caused by prior Socratic exposure rather than by who chose to keep using the chatbot, uneven team makeup, and only partial traces of who prompted how.","fun_headline_variants_meta":{"raw":{"variants":["Socratic tutors train lasting free-use prompting habits","Questioning tutors beat prompt tips for later LLM skills","Socratic guidance raises later gains and smarter free LLM use","Dialogic tutors build understanding-driven prompting after scaffolds","Socratic LLM scaffolding yields better independent prompting"]},"model":"grok-4.5","effort":"low","cost_usd":0.005072,"raw_usage":{"total_tokens":1420,"prompt_tokens":766,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":50720000,"prompt_tokens_details":{"text_tokens":766,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":577,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":766,"tokens_out":77,"duration_ms":4796,"temperature":1.0,"reasoning_tokens":577,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T03:23:43.865738+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A larger randomized follow-up that forces full individual-level tracing of unconstrained LLM use and individual (not group oral) understanding scores: if Socratic assignment no longer predicts understanding-driven clusters or understanding once self-selection and composition are blocked, the transfer claim fails.","supporting_citations":[],"review_version":1}