REVIEW 3 major objections 5 minor 26 references
General-purpose helpfulness scores cannot tell answer-giving tutors from pedagogical ones.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 16:58 UTC pith:AI5YBMQ6
load-bearing objection Careful pre-registered audit showing helpfulness is a weak sole pedagogy signal in a controlled tutor testbed; the dissociation is real but partly instrument-shaped, and the process measures are the durable part. the 3 major comments →
Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In a pre-registered, controlled audit across three tutor bases, general-purpose LLM-judged helpfulness is not a reliable pedagogy signal. On the primary base under the frozen primary judge, conversational and pedagogical policies do not differ significantly in helpfulness (Cliff’s |δ|=0.10) but are perfectly rank-separated under a pedagogy rubric (|δ|=1.0). Across two judges, helpfulness orderings reverse on two of three bases while pedagogy contrasts retain their direction where detected; seven policies span 2.3 points in judged pedagogy inside a 0.25-point helpfulness band; and answer-revealing turns are followed by less independent student work on every base.
What carries the argument
A within-base policy contrast that freezes model weights and varies only control flow and prompts (minimal ConvTutor vs routed PedTutor), scored by two frozen condition-blind LLM judges on the same answer-phase turns plus two deterministic detectors—answer leakage by string match and next-turn student independence by explicit reasoning steps—so judged helpfulness can be compared against both a pedagogy rubric and judge-invariant process.
Load-bearing premise
That one deliberately weak simulated student on a single algebra word-problem set is enough of an instrument for the failure of helpfulness to track pedagogy here to support treating helpfulness as unreliable for tutor evaluation in general.
What would settle it
Re-run the same frozen transcripts and policies with human raters (or additional independent judges) under the identical helpfulness and pedagogy rubrics: if helpfulness then cleanly and consistently ranks the pedagogical policy above the answer-giving one in the same direction as pedagogy and as the leakage–independence coupling, across bases without judge reversals, the central unreliability claim fails.
If this is right
- Tutor leaderboards and preference datasets that optimize only generic helpfulness can reward answer-giving rather than scaffolding.
- Evaluation pipelines should pair a pedagogy-targeted rubric with deterministic leakage and independence checks rather than a single helpfulness score.
- Process detectors (answer reveal; next-turn student work) can serve as filters or penalties when building pedagogical preference pairs.
- Judge choice alone can flip which tutoring policy looks “more helpful,” so single-judge helpfulness rankings of tutors are not stable.
Where Pith is reading between the lines
- Reward models trained on ordinary assistant helpfulness may systematically push tutoring systems toward disclosure unless pedagogy-specific preference data or process penalties are added.
- The same dissociation may appear in other “helpfulness vs. long-horizon goal” settings (coaching, debugging help, medical triage explanations) where in-the-moment usefulness diverges from skill-building.
- A cheap deployment check—flag turns that match known answer forms and measure whether the next user turn still shows work—could catch failure modes that scalar helpfulness misses.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper audits whether a general-purpose LLM-judged helpfulness rubric can distinguish answer-giving from pedagogical guidance in tutoring. Within each of three tutor bases, ConvTutor and PedTutor share identical frozen weights and face one weak simulated student; Claude Opus 4.8 is the frozen primary judge, with a prospectively specified GPT-5.6 Sol robustness audit on the same 1,179 answer-phase turns. Deterministic detectors track answer leakage and next-turn independent work. On the primary base under Opus, session helpfulness does not separate the policies (Cliff’s |δ|=0.10) while a pedagogy rubric yields perfect rank separation (|δ|=1.0); helpfulness orderings reverse between judges on two of three bases, whereas pedagogy direction is retained where detected. An ablation shows seven policies spanning 2.3 pedagogy points inside a 0.25 helpfulness band. Leakage is followed by less independent student work on every base. The authors conclude that general-purpose helpfulness is not a reliable sole pedagogy signal and recommend pairing pedagogy-targeted rubrics with process measures.
Significance. If the result holds in this controlled setting, it is a useful construct-validity warning for a rapidly growing practice: using generic helpfulness/preference judges to select or optimize LLM tutors. Strengths that raise the contribution above a null-result note include pre-registered decision rules (Table 1), identical-weight policy isolation, frozen judge and rubrics, deterministic judge-invariant process labels, cross-base replications, a matched visible-turn sensitivity, an explicit manipulation-check vs live-test separation, and released code. The judge-contingent helpfulness ordering and the leakage→independence coupling are the most portable findings. Scope is appropriately narrow (one weak simulator, one algebra domain, no human validation, no claim of durable learning), so the paper’s value is as a measurement audit rather than a general theory of tutoring quality.
major comments (3)
- [Abstract, §4, Table 1, Table 4] Abstract and §4 lead with the Opus primary-base contrast Cliff’s |δ|=0.10 (helpfulness, n.s.) vs |δ|=1.0 (pedagogy). Per §3.2 and Table 4, the pedagogy rubric was registered after primary helpfulness scores were frozen and scores the same four principles PedTutor encodes, so the perfect separation is owned as a manipulation check, not independent evidence of better teaching. The pre-registered live test was P2 (helpfulness Conv>Ped); J1 failed on every base (Table 1). The manuscript’s body is careful, but the abstract’s headline pairing still reads as a confirmatory double result. Reframe the abstract and opening of §4 so the primary confirmatory content is the failed/null or judge-contingent helpfulness signal, the turn-level J2 coupling, and the cross-judge reversal—with pedagogy |δ|=1.0 explicitly labeled a manipulation check in the same sentence.
- [Appendix E.1, §5, §7, J2 in §4] Appendix E.1’s helpfulness rubric explicitly instructs the judge not to reward or penalize disclosing vs withholding the answer, and not to judge teaching strategy. That design makes the instrument a fair test of whether a strategy-blind general-purpose signal still tracks pedagogy, but it also means a null or small session-level helpfulness gap is partly instrument-driven—especially under severe ceiling compression (Opus median 5 on 975/1179 turns; §5). Meanwhile J2 finds a positive leakage→helpfulness association on the primary base under Opus (+0.303). §7 should state more sharply how a strategy-excluding, ceilinged rubric can still show turn-level reward for leakage yet fail as a session-level ranking signal, so readers do not treat the |δ|=0.10 null as pure evidence that helpfulness “misses” answer-giving rather than as evidence it is an unreliable sole ranking criterion under this
- [§7, §3.1, §A.3–A.5, Abstract] The central generalization—“general-purpose helpfulness is not a reliable pedagogy signal”—rests on one deliberately weak student (Llama-3.1-8B, ~11% isolated accuracy), carried-context probes that saturate (§3.2, §A.3–A.5), n=10 paired replicates per cell, and no human annotator validation of either rubric. §7 already lists these boundaries; they are load-bearing for how far the evaluation recommendation travels. Add a short, concrete external-validity paragraph (or revise §7) stating which claim tiers are supported: (i) within this testbed; (ii) as a caution for LLM-as-judge tutor ranking without process measures; (iii) not licensed for human learning or for strong-student regimes. Avoid letting the abstract’s closing prescription outrun tier (i)–(ii).
minor comments (5)
- [Table 2, Figure 2, Appendix B.7] Figure 2 caption and Table 2 note that δ can disagree in sign with paired ∆ (Gemini/Sol pedagogy). Consider reporting the paired Hodges–Lehmann estimate alongside ∆ in Table 2a to reduce reader confusion when δ and ∆ diverge.
- [§5, Table 2] Gemini’s provider-enforced minimum reasoning setting is declared (§3.1, Table 2 footnote) but easy to miss when scanning cross-base claims. Flag it once in the main-text results paragraph that first pools “every base.”
- [§4, Appendix B.9] The matched-budget analysis leaves PedTutor with only six leaky turns (§4), below the authors’ own adequacy floor of ten. The text notes this; a single sentence in the Table 2 or §4 discussion that J2 under the matched window is ConvTutor-dominated would help skimmers.
- [Throughout] Minor prose/typo cleanup: spacing artifacts in the compiled text (e.g., “pre-registeredstudy”, “policiesdonotdiffersignificantly”) and inconsistent model-name hyphenation should be normalized in production.
- [§2] Related Work could briefly cite human ITS evaluation practice (e.g., knowledge-component or step-level process metrics) to situate why deterministic leakage/independence are natural complements to LLM judges.
Circularity Check
Disclosed construction match: pedagogy separation is a manipulation check on principles PedTutor was built from; helpfulness unreliability still has independent legs.
specific steps
-
self definitional
[§3.2; also §A.1, Abstract, Table 4]
"Because PedTutor is built on the same principles this rubric scores, any pedagogy gap between the tutors is reported as a manipulation check rather than independent evidence. The construction-independent signal is the divergence between judged helpfulness and judged pedagogy on the same turns (§4)."
The large pedagogy contrast (primary-base Opus |δ|=1.0, perfect replicate rank separation) is induced in part by scoring PedTutor against a rubric written to the same four principles used to construct PedTutor’s tracker/responders. Using that gap as one pole of the claim that helpfulness ‘misses pedagogical differences’ makes the dissociation partly definitional on the pedagogy pole. The paper correctly labels it a manipulation check, but the headline unreliability claim still consumes that constructed separation as contrast evidence.
-
self definitional
[App. E.1 helpfulness rubric; Abstract research question; §3.2]
"Do NOT reward or penalize any particular tutoring strategy. In particular, do not consider whether the tutor gave away the answer or withheld it, whether it pushed the student to keep trying, or whether it encouraged the student to reason independently. Those choices are outside this rating: score only clarity, responsiveness, and in-the-moment helpfulness."
The abstract asks whether a general-purpose helpfulness rubric can distinguish direct answer-giving from pedagogical guidance, while the frozen helpfulness instrument explicitly instructs the judge not to treat disclose-vs-withhold as in-scope. Session-level failure of helpfulness to separate leaky ConvTutor from withholding PedTutor is therefore partly aligned with rubric design (strategy-blindness plus ceiling compression), not a pure discovery that an unrestricted preference signal is blind. Mitigating: turn-level J2 still finds leaky turns rated more helpful under Opus, so the instrument is not fully inert to disclosure.
full rationale
This is an empirical audit, not a fitted theoretical derivation, so classical fit-as-prediction circularity is absent. The one clear construction dependence is owned in the text: PedTutor’s nodes implement the same four learning-sciences principles that the pedagogy rubric scores, so the perfect primary-base pedagogy rank separation (|δ|=1.0) is a manipulation check, not independent proof that PedTutor is better teaching. The paper’s load-bearing claim—that general-purpose helpfulness is not a reliable pedagogy signal—partly leans on the divergence between that by-construction pedagogy gap and a small/null helpfulness gap. That is moderate circularity on one arm of the contrast. It is not fatal to the paper: (i) the pre-registered live test was P2 (helpfulness Conv>Ped), which failed rather than being reverse-engineered; (ii) judge-contingent helpfulness orderings (Opus vs Sol reversals on two bases), the deterministic leakage→next-turn-independence coupling on every base, and the seven-policy helpfulness compression (0.25-point band vs 2.3 pedagogy) do not reduce to the PedTutor–rubric match; (iii) there is no self-citation uniqueness chain or renamed known theorem. Score 4 reflects disclosed, partial construction dependence on the pedagogy arm with independent content remaining in the central claim.
Axiom & Free-Parameter Ledger
free parameters (5)
- Session replicate count (n=10 pairs per condition per base) =
10 paired replicates
- Tutor sampling temperature and completion limits =
tutor T=0.4; student T=0.8
- Fixed four tutor turns per training problem alternation =
4 tutor turns per training problem
- PedTutor routing thresholds (hint level, ≥2 reasoning attempts before concrete hints) =
attempt threshold=2; hint level=prior tutor turns capped below concrete
- Judge repetition count and holistic 1–5 aggregation =
3 ratings; mean overall in {1..5}
axioms (7)
- domain assumption A general-purpose helpfulness rubric that explicitly ignores disclosure strategy is a fair test of whether annotator-style helpfulness detects pedagogical differences.
- domain assumption String-match leakage against pre-declared numeric/phrase forms is a conservative operational index of answer disclosure (paraphrases missed).
- domain assumption Symbolic/numeric work detectors on student turns validly index ‘independent work’ without model adjudication.
- domain assumption One weak frozen student simulator makes tutoring load-bearing enough that process contrasts are meaningful.
- standard math Wilcoxon signed-rank on 10 paired replicate differences and crossed mixed models for turn-level β_L are appropriate confirmatory tests at α=0.05.
- ad hoc to paper Pedagogy-rubric separation is reported only as a manipulation check because PedTutor encodes the scored principles.
- domain assumption LLM judges without human validation still bound robustness across evaluators for the reliability claim inside this setting.
invented entities (5)
-
ConvTutor vs PedTutor policy layer (identical base weights)
no independent evidence
-
Answer-leakage detector (token-boundary numeric + phrase lists)
independent evidence
-
Next-turn independence / session independence ratio
independent evidence
-
Symmetric judged-pedagogy rubric (four principles + holistic)
no independent evidence
-
Frozen 19-problem mixture/weighted-average instrument with isomorphs
independent evidence
read the original abstract
LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measure answer leakage and next-turn independent work. Claude Opus 4.8 is the frozen, condition-blind primary judge. After the Opus scores were fixed, GPT-5.6 Sol was prospectively specified for a post hoc robustness audit of the same 1,179 confirmatory answer-phase tutor turns under the frozen helpfulness and pedagogy rubrics. On the primary base under Opus, the policies do not differ significantly in helpfulness but are perfectly rank-separated under the pedagogy rubric (Cliff's $|\delta|{=}0.10$ vs. $1.0$). Across the two judges, pedagogy contrasts retain their direction where detected, whereas the helpfulness ordering is judge-contingent, reversing between judges on two of three bases. In an Opus-only ablation, seven primary-base policies span $2.3$ points in mean judged pedagogy within a $0.25$-point band of mean judged helpfulness. Separately, answer-revealing turns are followed by less independent student work on every base, a result that is judge-invariant by construction. In this controlled setting, general-purpose helpfulness is not a reliable pedagogy signal. Tutor evaluation should pair pedagogy-targeted rubrics with deterministic process measures.
Figures
Reference graph
Works this paper leans on
-
[1]
Advances in neural information processing systems , volume=
Training language models to follow instructions with human feedback , author=. Advances in neural information processing systems , volume=
-
[2]
Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=
Pedagogical alignment of large language models , author=. Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=
2024
-
[3]
Findings of the Association for Computational Linguistics: ACL 2025 , pages=
Towards the pedagogical steering of large language models for tutoring: A case study with modeling productive failure , author=. Findings of the Association for Computational Linguistics: ACL 2025 , pages=
2025
-
[4]
Proceedings of the National Academy of Sciences , volume=
Generative AI without guardrails can harm learning: Evidence from high school mathematics , author=. Proceedings of the National Academy of Sciences , volume=. 2025 , publisher=
2025
-
[5]
arXiv preprint arXiv:2410.21819 , year=
Self-preference bias in llm-as-a-judge , author=. arXiv preprint arXiv:2410.21819 , year=
-
[6]
Findings of the Association for Computational Linguistics: ACL 2024 , pages=
Benchmarking cognitive biases in large language models as evaluators , author=. Findings of the Association for Computational Linguistics: ACL 2024 , pages=
2024
-
[7]
Judging the judges: A systematic study of position bias in llm-as-a-judge , author=. Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics , pages=
-
[8]
International Conference on Learning Representations , volume=
Evaluating large language models at evaluating instruction following , author=. International Conference on Learning Representations , volume=
-
[9]
Findings of the Association for Computational Linguistics: NAACL 2025 , pages=
Rewardbench: Evaluating reward models for language modeling , author=. Findings of the Association for Computational Linguistics: NAACL 2025 , pages=
2025
-
[10]
Advances in Neural Information Processing Systems , volume=
SocraticLM: Exploring socratic personalized teaching with large language models , author=. Advances in Neural Information Processing Systems , volume=
-
[11]
Proceedings of the Eleventh ACM Conference on Learning@ Scale , pages=
Autotutor meets large language models: A language model tutor with rich pedagogy and guardrails , author=. Proceedings of the Eleventh ACM Conference on Learning@ Scale , pages=
-
[12]
, author=
The role of tutoring in problem solving. , author=. Journal of child psychology and psychiatry, and allied disciplines , year=
-
[13]
Journal of Experimental Psychology: Human Learning & Memory , year=
The Generation Effect: Delineation of a Phenomenon , author=. Journal of Experimental Psychology: Human Learning & Memory , year=
-
[14]
Educational psychology review , volume=
Exploring the assistance dilemma in experiments with cognitive tutors , author=. Educational psychology review , volume=. 2007 , publisher=
2007
-
[15]
Toward Meta-cognitive Tutoring: A Model of Help Seeking with a Cognitive Tutor , author=. Int. J. Artif. Intell. Educ. , year=
-
[16]
arXiv preprint arXiv:2601.05473 , year=
Towards Valid Student Simulation with Large Language Models , author=. arXiv preprint arXiv:2601.05473 , year=
-
[17]
arXiv preprint arXiv:2601.04025 , year=
Simulated Students in Tutoring Dialogues: Substance or Illusion? , author=. arXiv preprint arXiv:2601.04025 , year=
-
[18]
2023 , eprint=
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena , author=. 2023 , eprint=
2023
-
[19]
International Conference on Learning Representations , volume=
Towards understanding sycophancy in language models , author=. International Conference on Learning Representations , volume=
-
[20]
2025 , eprint=
Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach , author=. 2025 , eprint=
2025
-
[21]
2025 , eprint=
LearnLM: Improving Gemini for Learning , author=. 2025 , eprint=
2025
-
[22]
2024 , eprint=
The Llama 3 Herd of Models , author=. 2024 , eprint=
2024
-
[23]
2026 , month = feb, howpublished =
2026
-
[24]
2026 , month = may, howpublished =
2026
-
[25]
2026 , month = apr, howpublished =
2026
-
[26]
2026 , month = jul, howpublished =
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.