Pith. sign in

REVIEW 3 major objections 5 minor 26 references

General-purpose helpfulness scores cannot tell answer-giving tutors from pedagogical ones.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 16:58 UTC pith:AI5YBMQ6

load-bearing objection Careful pre-registered audit showing helpfulness is a weak sole pedagogy signal in a controlled tutor testbed; the dissociation is real but partly instrument-shaped, and the process measures are the durable part. the 3 major comments →

arxiv 2607.28128 v1 pith:AI5YBMQ6 submitted 2026-07-30 cs.CL cs.AIcs.CY

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

classification cs.CL cs.AIcs.CY
keywords LLM-as-judgetutoring evaluationhelpfulness rubricpedagogical alignmentanswer leakageprocess measuressimulated studentspre-registered audit
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

When large language models tutor students, evaluators often score them with the same generic “helpfulness” rubrics used for ordinary assistants. This paper asks whether that signal can tell a tutor that hands over the answer from one that scaffolds the student to work it out. Holding the underlying model fixed and swapping only the tutoring policy, the authors find that a frozen helpfulness judge barely separates the two policies, while a pedagogy-targeted rubric separates them perfectly. A second judge even reverses which policy looks more helpful on two of three model bases, yet keeps the pedagogy ranking. Separately, turns that leak the answer are followed by less independent student work on every base—a process fact no judge can rewrite. The upshot is practical: ranking tutors by generic helpfulness alone is an unreliable way to prefer pedagogical guidance over answer-giving.

Core claim

In a pre-registered, controlled audit across three tutor bases, general-purpose LLM-judged helpfulness is not a reliable pedagogy signal. On the primary base under the frozen primary judge, conversational and pedagogical policies do not differ significantly in helpfulness (Cliff’s |δ|=0.10) but are perfectly rank-separated under a pedagogy rubric (|δ|=1.0). Across two judges, helpfulness orderings reverse on two of three bases while pedagogy contrasts retain their direction where detected; seven policies span 2.3 points in judged pedagogy inside a 0.25-point helpfulness band; and answer-revealing turns are followed by less independent student work on every base.

What carries the argument

A within-base policy contrast that freezes model weights and varies only control flow and prompts (minimal ConvTutor vs routed PedTutor), scored by two frozen condition-blind LLM judges on the same answer-phase turns plus two deterministic detectors—answer leakage by string match and next-turn student independence by explicit reasoning steps—so judged helpfulness can be compared against both a pedagogy rubric and judge-invariant process.

Load-bearing premise

That one deliberately weak simulated student on a single algebra word-problem set is enough of an instrument for the failure of helpfulness to track pedagogy here to support treating helpfulness as unreliable for tutor evaluation in general.

What would settle it

Re-run the same frozen transcripts and policies with human raters (or additional independent judges) under the identical helpfulness and pedagogy rubrics: if helpfulness then cleanly and consistently ranks the pedagogical policy above the answer-giving one in the same direction as pedagogy and as the leakage–independence coupling, across bases without judge reversals, the central unreliability claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Tutor leaderboards and preference datasets that optimize only generic helpfulness can reward answer-giving rather than scaffolding.
  • Evaluation pipelines should pair a pedagogy-targeted rubric with deterministic leakage and independence checks rather than a single helpfulness score.
  • Process detectors (answer reveal; next-turn student work) can serve as filters or penalties when building pedagogical preference pairs.
  • Judge choice alone can flip which tutoring policy looks “more helpful,” so single-judge helpfulness rankings of tutors are not stable.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Reward models trained on ordinary assistant helpfulness may systematically push tutoring systems toward disclosure unless pedagogy-specific preference data or process penalties are added.
  • The same dissociation may appear in other “helpfulness vs. long-horizon goal” settings (coaching, debugging help, medical triage explanations) where in-the-moment usefulness diverges from skill-building.
  • A cheap deployment check—flag turns that match known answer forms and measure whether the next user turn still shows work—could catch failure modes that scalar helpfulness misses.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper audits whether a general-purpose LLM-judged helpfulness rubric can distinguish answer-giving from pedagogical guidance in tutoring. Within each of three tutor bases, ConvTutor and PedTutor share identical frozen weights and face one weak simulated student; Claude Opus 4.8 is the frozen primary judge, with a prospectively specified GPT-5.6 Sol robustness audit on the same 1,179 answer-phase turns. Deterministic detectors track answer leakage and next-turn independent work. On the primary base under Opus, session helpfulness does not separate the policies (Cliff’s |δ|=0.10) while a pedagogy rubric yields perfect rank separation (|δ|=1.0); helpfulness orderings reverse between judges on two of three bases, whereas pedagogy direction is retained where detected. An ablation shows seven policies spanning 2.3 pedagogy points inside a 0.25 helpfulness band. Leakage is followed by less independent student work on every base. The authors conclude that general-purpose helpfulness is not a reliable sole pedagogy signal and recommend pairing pedagogy-targeted rubrics with process measures.

Significance. If the result holds in this controlled setting, it is a useful construct-validity warning for a rapidly growing practice: using generic helpfulness/preference judges to select or optimize LLM tutors. Strengths that raise the contribution above a null-result note include pre-registered decision rules (Table 1), identical-weight policy isolation, frozen judge and rubrics, deterministic judge-invariant process labels, cross-base replications, a matched visible-turn sensitivity, an explicit manipulation-check vs live-test separation, and released code. The judge-contingent helpfulness ordering and the leakage→independence coupling are the most portable findings. Scope is appropriately narrow (one weak simulator, one algebra domain, no human validation, no claim of durable learning), so the paper’s value is as a measurement audit rather than a general theory of tutoring quality.

major comments (3)
  1. [Abstract, §4, Table 1, Table 4] Abstract and §4 lead with the Opus primary-base contrast Cliff’s |δ|=0.10 (helpfulness, n.s.) vs |δ|=1.0 (pedagogy). Per §3.2 and Table 4, the pedagogy rubric was registered after primary helpfulness scores were frozen and scores the same four principles PedTutor encodes, so the perfect separation is owned as a manipulation check, not independent evidence of better teaching. The pre-registered live test was P2 (helpfulness Conv>Ped); J1 failed on every base (Table 1). The manuscript’s body is careful, but the abstract’s headline pairing still reads as a confirmatory double result. Reframe the abstract and opening of §4 so the primary confirmatory content is the failed/null or judge-contingent helpfulness signal, the turn-level J2 coupling, and the cross-judge reversal—with pedagogy |δ|=1.0 explicitly labeled a manipulation check in the same sentence.
  2. [Appendix E.1, §5, §7, J2 in §4] Appendix E.1’s helpfulness rubric explicitly instructs the judge not to reward or penalize disclosing vs withholding the answer, and not to judge teaching strategy. That design makes the instrument a fair test of whether a strategy-blind general-purpose signal still tracks pedagogy, but it also means a null or small session-level helpfulness gap is partly instrument-driven—especially under severe ceiling compression (Opus median 5 on 975/1179 turns; §5). Meanwhile J2 finds a positive leakage→helpfulness association on the primary base under Opus (+0.303). §7 should state more sharply how a strategy-excluding, ceilinged rubric can still show turn-level reward for leakage yet fail as a session-level ranking signal, so readers do not treat the |δ|=0.10 null as pure evidence that helpfulness “misses” answer-giving rather than as evidence it is an unreliable sole ranking criterion under this
  3. [§7, §3.1, §A.3–A.5, Abstract] The central generalization—“general-purpose helpfulness is not a reliable pedagogy signal”—rests on one deliberately weak student (Llama-3.1-8B, ~11% isolated accuracy), carried-context probes that saturate (§3.2, §A.3–A.5), n=10 paired replicates per cell, and no human annotator validation of either rubric. §7 already lists these boundaries; they are load-bearing for how far the evaluation recommendation travels. Add a short, concrete external-validity paragraph (or revise §7) stating which claim tiers are supported: (i) within this testbed; (ii) as a caution for LLM-as-judge tutor ranking without process measures; (iii) not licensed for human learning or for strong-student regimes. Avoid letting the abstract’s closing prescription outrun tier (i)–(ii).
minor comments (5)
  1. [Table 2, Figure 2, Appendix B.7] Figure 2 caption and Table 2 note that δ can disagree in sign with paired ∆ (Gemini/Sol pedagogy). Consider reporting the paired Hodges–Lehmann estimate alongside ∆ in Table 2a to reduce reader confusion when δ and ∆ diverge.
  2. [§5, Table 2] Gemini’s provider-enforced minimum reasoning setting is declared (§3.1, Table 2 footnote) but easy to miss when scanning cross-base claims. Flag it once in the main-text results paragraph that first pools “every base.”
  3. [§4, Appendix B.9] The matched-budget analysis leaves PedTutor with only six leaky turns (§4), below the authors’ own adequacy floor of ten. The text notes this; a single sentence in the Table 2 or §4 discussion that J2 under the matched window is ConvTutor-dominated would help skimmers.
  4. [Throughout] Minor prose/typo cleanup: spacing artifacts in the compiled text (e.g., “pre-registeredstudy”, “policiesdonotdiffersignificantly”) and inconsistent model-name hyphenation should be normalized in production.
  5. [§2] Related Work could briefly cite human ITS evaluation practice (e.g., knowledge-component or step-level process metrics) to situate why deterministic leakage/independence are natural complements to LLM judges.

Circularity Check

2 steps flagged

Disclosed construction match: pedagogy separation is a manipulation check on principles PedTutor was built from; helpfulness unreliability still has independent legs.

specific steps
  1. self definitional [§3.2; also §A.1, Abstract, Table 4]
    "Because PedTutor is built on the same principles this rubric scores, any pedagogy gap between the tutors is reported as a manipulation check rather than independent evidence. The construction-independent signal is the divergence between judged helpfulness and judged pedagogy on the same turns (§4)."

    The large pedagogy contrast (primary-base Opus |δ|=1.0, perfect replicate rank separation) is induced in part by scoring PedTutor against a rubric written to the same four principles used to construct PedTutor’s tracker/responders. Using that gap as one pole of the claim that helpfulness ‘misses pedagogical differences’ makes the dissociation partly definitional on the pedagogy pole. The paper correctly labels it a manipulation check, but the headline unreliability claim still consumes that constructed separation as contrast evidence.

  2. self definitional [App. E.1 helpfulness rubric; Abstract research question; §3.2]
    "Do NOT reward or penalize any particular tutoring strategy. In particular, do not consider whether the tutor gave away the answer or withheld it, whether it pushed the student to keep trying, or whether it encouraged the student to reason independently. Those choices are outside this rating: score only clarity, responsiveness, and in-the-moment helpfulness."

    The abstract asks whether a general-purpose helpfulness rubric can distinguish direct answer-giving from pedagogical guidance, while the frozen helpfulness instrument explicitly instructs the judge not to treat disclose-vs-withhold as in-scope. Session-level failure of helpfulness to separate leaky ConvTutor from withholding PedTutor is therefore partly aligned with rubric design (strategy-blindness plus ceiling compression), not a pure discovery that an unrestricted preference signal is blind. Mitigating: turn-level J2 still finds leaky turns rated more helpful under Opus, so the instrument is not fully inert to disclosure.

full rationale

This is an empirical audit, not a fitted theoretical derivation, so classical fit-as-prediction circularity is absent. The one clear construction dependence is owned in the text: PedTutor’s nodes implement the same four learning-sciences principles that the pedagogy rubric scores, so the perfect primary-base pedagogy rank separation (|δ|=1.0) is a manipulation check, not independent proof that PedTutor is better teaching. The paper’s load-bearing claim—that general-purpose helpfulness is not a reliable pedagogy signal—partly leans on the divergence between that by-construction pedagogy gap and a small/null helpfulness gap. That is moderate circularity on one arm of the contrast. It is not fatal to the paper: (i) the pre-registered live test was P2 (helpfulness Conv>Ped), which failed rather than being reverse-engineered; (ii) judge-contingent helpfulness orderings (Opus vs Sol reversals on two bases), the deterministic leakage→next-turn-independence coupling on every base, and the seven-policy helpfulness compression (0.25-point band vs 2.3 pedagogy) do not reduce to the PedTutor–rubric match; (iii) there is no self-citation uniqueness chain or renamed known theorem. Score 4 reflects disclosed, partial construction dependence on the pedagogy arm with independent content remaining in the central claim.

Axiom & Free-Parameter Ledger

5 free parameters · 7 axioms · 5 invented entities

The central claim is empirical and rests on design choices and measurement definitions rather than fitted physical constants. Load-bearing assumptions are the weak simulated student as instrument, string-match leakage as a lower bound on disclosure, rule-based independence labels, condition-blind LLM judges as proxies for annotator preference (not human-validated), and the answer-phase window. Invented entities are operational constructs (policies, detectors, rubrics), not new physical objects.

free parameters (5)
  • Session replicate count (n=10 pairs per condition per base) = 10 paired replicates
    Fixed sample size that sets power; paper notes nulls only rule out large effects (d≳1).
  • Tutor sampling temperature and completion limits = tutor T=0.4; student T=0.8
    Serving configuration chosen for the study (temp 0.4, 1024-token cap; student temp 0.8); affects realized leakage and dialogue length.
  • Fixed four tutor turns per training problem alternation = 4 tutor turns per training problem
    Protocol length fixed in implementation before confirmatory collection; shapes visible-turn budgets and windows.
  • PedTutor routing thresholds (hint level, ≥2 reasoning attempts before concrete hints) = attempt threshold=2; hint level=prior tutor turns capped below concrete
    Hand-specified control-flow parameters that induce the withholding policy being audited.
  • Judge repetition count and holistic 1–5 aggregation = 3 ratings; mean overall in {1..5}
    Three ratings per turn per rubric; mean holistic score is the metric—design choice affecting variance and ceiling.
axioms (7)
  • domain assumption A general-purpose helpfulness rubric that explicitly ignores disclosure strategy is a fair test of whether annotator-style helpfulness detects pedagogical differences.
    Stated in §3.2 and Appendix E.1; underpins interpreting P2/J2 helpfulness legs as construct-validity evidence.
  • domain assumption String-match leakage against pre-declared numeric/phrase forms is a conservative operational index of answer disclosure (paraphrases missed).
    Appendix B.2; policy contrasts transfer to total disclosure only if undetected paraphrase rates are similar.
  • domain assumption Symbolic/numeric work detectors on student turns validly index ‘independent work’ without model adjudication.
    Appendix B.3; registered model confirmation pass was not run on confirmatory data (disclosed deviation).
  • domain assumption One weak frozen student simulator makes tutoring load-bearing enough that process contrasts are meaningful.
    §3.1, §A.5; selection gates fixed pre-confirmatory collection.
  • standard math Wilcoxon signed-rank on 10 paired replicate differences and crossed mixed models for turn-level β_L are appropriate confirmatory tests at α=0.05.
    §3.3, Appendix B.7–B.8; standard nonparametrics and linear mixed/LPM specs.
  • ad hoc to paper Pedagogy-rubric separation is reported only as a manipulation check because PedTutor encodes the scored principles.
    §3.2; prevents treating pedagogy δ=1.0 as independent evidence of superior teaching.
  • domain assumption LLM judges without human validation still bound robustness across evaluators for the reliability claim inside this setting.
    §2, §7; paper does not claim human preference validity.
invented entities (5)
  • ConvTutor vs PedTutor policy layer (identical base weights) no independent evidence
    purpose: Isolate tutoring policy (emergent answer-giving vs routed scaffolding) as the sole manipulation.
    Core experimental construct; not a physical entity but the treatment definition.
  • Answer-leakage detector (token-boundary numeric + phrase lists) independent evidence
    purpose: Deterministic, judge-invariant measure of explicit solution disclosure.
    Operational instrument with known under-call of paraphrases; forms frozen pre-confirmatory data.
  • Next-turn independence / session independence ratio independent evidence
    purpose: Process measure of whether the student attempts reasoning after tutor turns.
    Rule-based labels aligned to generation-effect / productive-struggle ideas; judge-invariant by construction.
  • Symmetric judged-pedagogy rubric (four principles + holistic) no independent evidence
    purpose: Provide a pedagogy-targeted contrast to general helpfulness on the same turns.
    Pre-registered before pedagogy scoring but aligned to PedTutor’s design principles—manipulation-check status disclosed.
  • Frozen 19-problem mixture/weighted-average instrument with isomorphs independent evidence
    purpose: Controlled algebra domain isolating equation setup rather than arithmetic.
    Authored and constraint-checked before confirmatory run; enables leakage forms and probe structure.

pith-pipeline@v1.2.0-daily-grok45 · 35575 in / 4507 out tokens · 86962 ms · 2026-07-31T16:58:54.014953+00:00 · methodology

0 comments
read the original abstract

LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measure answer leakage and next-turn independent work. Claude Opus 4.8 is the frozen, condition-blind primary judge. After the Opus scores were fixed, GPT-5.6 Sol was prospectively specified for a post hoc robustness audit of the same 1,179 confirmatory answer-phase tutor turns under the frozen helpfulness and pedagogy rubrics. On the primary base under Opus, the policies do not differ significantly in helpfulness but are perfectly rank-separated under the pedagogy rubric (Cliff's $|\delta|{=}0.10$ vs. $1.0$). Across the two judges, pedagogy contrasts retain their direction where detected, whereas the helpfulness ordering is judge-contingent, reversing between judges on two of three bases. In an Opus-only ablation, seven primary-base policies span $2.3$ points in mean judged pedagogy within a $0.25$-point band of mean judged helpfulness. Separately, answer-revealing turns are followed by less independent student work on every base, a result that is judge-invariant by construction. In this controlled setting, general-purpose helpfulness is not a reliable pedagogy signal. Tutor evaluation should pair pedagogy-targeted rubrics with deterministic process measures.

Figures

Figures reproduced from arXiv: 2607.28128 by Boyuan Deng, Chongyang Gao, Hongyang Zhang, Jiale Liu, Mengyu Xu, Qiaoxin Yang, Shuyi Fan.

Figure 1
Figure 1. Figure 1: Study design and matched interaction excerpts. (a) Within each tutor base, ConvTutor and PedTutor use identical [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: (a) Means of ten paired session-level ConvTutor−PedTutor differences in judged helpfulness (left) and judged pedagogy (right) for each tutor base under Claude Opus 4.8 (filled squares) and GPT-5.6 Sol (open circles); faded markers denote p > .05 in the two-sided paired Wilcoxon test. In the helpfulness panel, the mean ordering reverses between judges on Sonnet and the GPT-5.5 base, although the Opus Sonnet… view at source ↗
Figure 3
Figure 3. Figure 3: Answer leakage versus next-turn student independence (answer-phase window): a deterministic outcome that Sol [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Design, protocol, and measurement window. (a) Both tutoring policies are instantiated from identical frozen weights, [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

26 extracted references · 2 linked inside Pith

  1. [1]

    Advances in neural information processing systems , volume=

    Training language models to follow instructions with human feedback , author=. Advances in neural information processing systems , volume=

  2. [2]

    Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

    Pedagogical alignment of large language models , author=. Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

  3. [3]

    Findings of the Association for Computational Linguistics: ACL 2025 , pages=

    Towards the pedagogical steering of large language models for tutoring: A case study with modeling productive failure , author=. Findings of the Association for Computational Linguistics: ACL 2025 , pages=

  4. [4]

    Proceedings of the National Academy of Sciences , volume=

    Generative AI without guardrails can harm learning: Evidence from high school mathematics , author=. Proceedings of the National Academy of Sciences , volume=. 2025 , publisher=

  5. [5]

    arXiv preprint arXiv:2410.21819 , year=

    Self-preference bias in llm-as-a-judge , author=. arXiv preprint arXiv:2410.21819 , year=

  6. [6]

    Findings of the Association for Computational Linguistics: ACL 2024 , pages=

    Benchmarking cognitive biases in large language models as evaluators , author=. Findings of the Association for Computational Linguistics: ACL 2024 , pages=

  7. [7]

    Judging the judges: A systematic study of position bias in llm-as-a-judge , author=. Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics , pages=

  8. [8]

    International Conference on Learning Representations , volume=

    Evaluating large language models at evaluating instruction following , author=. International Conference on Learning Representations , volume=

  9. [9]

    Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

    Rewardbench: Evaluating reward models for language modeling , author=. Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

  10. [10]

    Advances in Neural Information Processing Systems , volume=

    SocraticLM: Exploring socratic personalized teaching with large language models , author=. Advances in Neural Information Processing Systems , volume=

  11. [11]

    Proceedings of the Eleventh ACM Conference on Learning@ Scale , pages=

    Autotutor meets large language models: A language model tutor with rich pedagogy and guardrails , author=. Proceedings of the Eleventh ACM Conference on Learning@ Scale , pages=

  12. [12]

    , author=

    The role of tutoring in problem solving. , author=. Journal of child psychology and psychiatry, and allied disciplines , year=

  13. [13]

    Journal of Experimental Psychology: Human Learning & Memory , year=

    The Generation Effect: Delineation of a Phenomenon , author=. Journal of Experimental Psychology: Human Learning & Memory , year=

  14. [14]

    Educational psychology review , volume=

    Exploring the assistance dilemma in experiments with cognitive tutors , author=. Educational psychology review , volume=. 2007 , publisher=

  15. [15]

    Toward Meta-cognitive Tutoring: A Model of Help Seeking with a Cognitive Tutor , author=. Int. J. Artif. Intell. Educ. , year=

  16. [16]

    arXiv preprint arXiv:2601.05473 , year=

    Towards Valid Student Simulation with Large Language Models , author=. arXiv preprint arXiv:2601.05473 , year=

  17. [17]

    arXiv preprint arXiv:2601.04025 , year=

    Simulated Students in Tutoring Dialogues: Substance or Illusion? , author=. arXiv preprint arXiv:2601.04025 , year=

  18. [18]

    2023 , eprint=

    Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena , author=. 2023 , eprint=

  19. [19]

    International Conference on Learning Representations , volume=

    Towards understanding sycophancy in language models , author=. International Conference on Learning Representations , volume=

  20. [20]

    2025 , eprint=

    Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach , author=. 2025 , eprint=

  21. [21]

    2025 , eprint=

    LearnLM: Improving Gemini for Learning , author=. 2025 , eprint=

  22. [22]

    2024 , eprint=

    The Llama 3 Herd of Models , author=. 2024 , eprint=

  23. [23]

    2026 , month = feb, howpublished =

  24. [24]

    2026 , month = may, howpublished =

  25. [25]

    2026 , month = apr, howpublished =

  26. [26]

    2026 , month = jul, howpublished =