{"id":"a585abf0-afdc-4479-9eb5-e6a95fef0794","arxiv_id":"2607.28128","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"General-purpose LLM-judged helpfulness is not a reliable signal of whether a tutor is teaching versus giving answers.","lead":"A study of LLM math tutors found that a general-purpose helpfulness score did not reliably distinguish tutors that hand students the answer from tutors that guide them, and the ranking flipped depending on which AI judge was used. The paper argues tutor evaluation should pair pedagogy-specific rubrics with deterministic behavioral measures, like whether the answer was leaked.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Helpfulness rubric's explicit ban on considering answer disclosure narrows the central claim's scope","rationale":"The reader identified the same load-bearing concern: the helpfulness rubric explicitly excludes answer disclosure/withholding, making the central claim narrower than its 'general-purpose helpfulness' wording. I agree that this is the most significant soft spot. The paper is otherwise careful and well-controlled: it fixes base weights, pre-registers confirmatory claims, uses deterministic process measures, reports manipulation-check status of the pedagogy contrast, and transparently discloses specification timing. The cross-judge reversal and the deterministic leakage–independence coupling provide real evidence that the specific Rubric E.1 is not a robust pedagogy signal. However, the central claim is phrased as a statement about 'general-purpose helpfulness' generally, and the only helpfulness instrument tested is one that deliberately instructs the judge to ignore tutoring strategy, including answer-giving. This makes the finding less informative about how helpfulness evaluations are actually used in practice, where disclosure may be implicitly or explicitly relevant. A concrete replication with a standard, non-restricted helpfulness rubric would settle whether the conclusion generalizes. Since the reader's verdict is already CONDITIONAL and my concern is the same one, I do not recommend changing the verdict. The claim should remain conditional on the specific operationalization and on the two judges tested.","tokens_in":30489,"tokens_out":6812,"duration_ms":74895,"concrete_test":"Re-score the same 1,179 frozen answer-phase tutor turns with the same two judges (Opus 4.8 and GPT-5.6 Sol) under a standard helpfulness rubric that contains no instruction forbidding consideration of answer disclosure or withholding—for example, the generic MT-Bench helpfulness prompt, or a fixed reward-model API. Use the same three-repetition aggregation, answer-phase window, and policy-contrast analysis. If the helpfulness ordering becomes consistent across judges or separates ConvTutor and PedTutor in the pedagogy-consistent direction, the paper's conclusion must be rescoped to 'a strategy-blind helpfulness rubric is not a reliable pedagogy signal.' If the ordering still flips between judges or fails to separate, the current central claim survives this threat.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Appendix E.1's helpfulness rubric instructs the judge: 'Do NOT reward or penalize any particular tutoring strategy. In particular, do not consider whether the tutor gave away the answer or withheld it, whether it pushed the student to keep trying, or whether it encouraged the student to reason independently.' This operationalizes 'general-purpose helpfulness' as clarity + responsiveness + in-the-moment progress, deliberately excluding the dimension (answer disclosure) most closely tied to pedagogical quality. The paper defends this in §B.4 as a 'fair test,' arguing that a helpfulness instrument that already encoded a preference about answer-giving could not test whether helpfulness detects answer-giving. That defense conflates construct-validity with representativeness. Typical LLM-as-a-judge helpfulness rubrics (e.g., MT-Bench) and reward-model preference scores do not contain this instruction; they are allowed to treat answer-giving as helpful or unhelpful depending on context. If such a standard rubric separates the policies or remains stable across judges, the headline 'general-purpose helpfulness is not a reliable pedagogy signal' would be false in the intended general sense—it would hold only for the paper's strategy-blind variant. The cross-judge reversal and the leakage–independence coupling are real, but they are observed under this restricted rubric. Thus the evidence does not establish unreliability of general-purpose helpfulness as usually operationalized, only of a helpfulness construct that is deliberately blind to the very behavior (answer disclosure) being audited.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a pre-registered audit of whether LLM-judged helpfulness can distinguish direct answer-giving from pedagogical guidance in tutoring. Within each of three tutor base models, the authors compare a minimal conversational tutor (ConvTutor) with a structured pedagogical tutor (PedTutor), both instantiated from identical frozen weights and paired with one weak simulated student. They measure answer leakage and next-turn independent work deterministically, and they score the same answer-phase tutor turns with Claude Opus 4.8 (primary judge) and GPT-5.6 Sol (prospectively specified post hoc judge) under a helpfulness rubric and a pedagogy rubric. The main findings are: on the primary base, helpfulness does not separate the policies while judged pedagogy separates them perfectly; the helpfulness ordering reverses between judges on two of three bases; and leaky turns are followed by less independent student work on every base. The paper concludes that general-purpose helpfulness is not a reliable pedagogy signal in this controlled setting and recommends pairing pedagogy-targeted rubrics with deterministic process measures.","tokens_in":30766,"tokens_out":4792,"duration_ms":56787,"significance":"If the central claim were established, the paper would be a valuable cautionary result for LLM-as-a-judge evaluation of tutors. The study has genuine methodological strengths: it is pre-registered with a specification-status table (Appendix C), the rubrics and prompts are released verbatim, the primary judge is condition-blind, the process measures are deterministic and judge-invariant, sensitivity analyses (matched visible-turn budget, policy-adjusted specifications, clustered estimates) are reported, and the authors disclose deviations and post hoc components rather than hiding them. The cleanest result — that leaky turns are followed by less independent student work, across all bases and all seven policies — is robust and does not depend on any judge. However, the headline conclusion about 'general-purpose helpfulness' is narrower than the operationalization actually used, and the pedagogy criterion itself lacks external validation. These limitations bear directly on the paper's main claim.","major_comments":[{"comment":"The headline claim that 'general-purpose helpfulness is not a reliable pedagogy signal' overstates what the design tests. The helpfulness rubric used in this study explicitly instructs the judge: 'Do NOT consider whether the tutor gave away the answer or withheld it...' (Appendix E.1), and §B.4 frames this exclusion as necessary for a 'fair test.' But this is a strategy-blind helpfulness instrument, not the construct used in typical helpfulness evaluations such as MT-Bench-style rubrics or reward-model preference scores, which are free to weigh answer-giving when it is contextually relevant. The evidence therefore supports only the narrower conclusion that a helpfulness rubric which is deliberately blind to answer disclosure fails to detect the pedagogy contrast that a pedagogy rubric detects. To support the general claim in §8 and the abstract, the authors would need either to rescope t","section":"Abstract, §8, Appendix E.1 and §B.4"},{"comment":"The pedagogy contrast is explicitly a manipulation check because PedTutor is constructed from the same principles that the pedagogy rubric scores (§3.2). Consequently, the central divergence in §4 — a near-zero helpfulness contrast beside a maximal pedagogy contrast (|δ|=0.10 vs 1.0) — is a comparison between a strategy-blind rubric and a rubric that is aligned with one policy by construction. This does not establish that the helpfulness signal is unreliable with respect to pedagogical quality as an external construct. No human ratings, expert annotations, or independently validated pedagogical quality measures are provided to anchor the pedagogy rubric. The paper acknowledges this in §7, but the abstract and conclusion present the finding as a general failure of helpfulness. A small human-validation study, or at minimum a rephrasing that limits the conclusion to 'LLM-judged pedagogy' ve","section":"§3.2, §4, Appendix E.2"},{"comment":"The cross-judge 'reversal' on the Sonnet base is a sign flip between a nonsignificant Opus contrast (Δ=-0.094, p=.160) and a significant Sol contrast (Δ=+0.115, p=.0039). Calling this a reversal is descriptively fair, but with ten paired sessions per base the study has limited power to detect moderate effects, so the Sonnet pattern may reflect noise rather than a genuine judge-by-policy interaction. The GPT-5.5 base, where the two contrasts are both significant and opposite, is the only strong evidence of a true reversal. The discussion in §5 is careful about this ('ordering reversal, not a significant reversal'), but the abstract, which says the ordering 'reversing between judges on two of three bases,' loses that nuance. I would ask the authors to qualify the cross-judge claim in the abstract and to report the significance asymmetry more prominently.","section":"§5, Table 2a"}],"minor_comments":[{"comment":"Consider changing 'reversing between judges on two of three bases' to 'reversing sign on two of three bases, with a significant opposite ordering on one base,' to reflect the low session-level power.","section":"Abstract"},{"comment":"The caption defines filled squares and open circles, but it may help to restate the judge names in the caption itself rather than only in the main text.","section":"§5, Figure 2 caption"},{"comment":"The version-dependent tie correction for the Wilcoxon test is disclosed; a sentence giving the library and versions used for both the primary and replication bases would improve reproducibility.","section":"Appendix B.7"},{"comment":"The finding that the final-answer-ban variant still leaks on 14.6% of in-window turns is striking; a one-sentence comment on why the ban fails despite its explicit instruction would make the ablation more informative.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually transparent and the deterministic process-measure result is solid, but the central claim as worded is broader than the operationalization. The main fixes are rescoping the construct (or adding a standard-rubric robustness check) and tempering the conclusion given the absence of human validation. I do not see this as rejectable, since the core design and the leakage–independence coupling are valuable and could support a carefully qualified version of the thesis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The paper's real result is a clean demonstration that a helpfulness-only LLM judge can be blind to answer-giving: under a frozen judge, the two tutor policies were statistically indistinguishable on helpfulness while perfectly separated on a pedagogy rubric, and the helpfulness ordering reversed across a second judge on two of three bases. The deterministic leakage-to-independence coupling is the strongest part — judge-invariant and consistent on every base. Second, the \"general-purpose helpfulness\" in the abstract is not the general-purpose helpfulness most people use. Appendix E.1's rubric tells the judge not to consider whether the tutor gave away the answer. The stress-test note is right: MT-Bench-style rubrics and reward models do not carry that prohibition, so the finding holds for a strategy-blind helpfulness construct, not necessarily for helpfulness as usually operationalized. The paper's defense in B.4 — that a rubric with a preference about disclosure could not test whether helpfulness detects disclosure — is a genuine construct-validity argument, but it narrows the scope rather than removing the concern.\n\nWhat's good: the design is unusually transparent. Pre-registration, frozen decision rules, deterministic detectors defined before data, an explicit specification-status table, and honest reporting of post-hoc fixes (matched-turn budget, Sol audit). The seven-policy ablation and the compression of helpfulness into a 0.25-point band is a nice addition. No human annotation is a real limit, but it is disclosed and the claims don't overreach into human learning.\n\nThe soft spots are real but proportionate. n=10 paired sessions per base is low; the session-level tests only rule out large effects. The judge-tutor family overlap is a confound the authors acknowledge. The matched-budget rule was fixed after reading full-window results — disclosed, but still a robustness check rather than confirmatory. The pedagogy separation is a manipulation check by construction. None of these sinks the paper; the central dissociation rests on effect-size contrast and the deterministic coupling, not on a single null.\n\nBottom line: this deserves a serious referee. The right fix is to rewrite the abstract and conclusion to say \"a strategy-blind helpfulness rubric\" rather than \"general-purpose helpfulness,\" and to add a sentence that typical helpfulness instruments without the disclosure-exclusion may behave differently. That narrows the contribution but makes it accurate.","headline":"A scrupulously honest audit showing helpfulness rubrics can miss answer-giving, but the headline overclaims generality because the rubric explicitly ignores disclosure.","tokens_in":31242,"tokens_out":1893,"would_cite":true,"duration_ms":21795,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"General-purpose helpfulness ratings from LLM judges cannot reliably distinguish tutors that give answers away from tutors that teach.","keywords":["LLM-as-a-judge","pedagogical alignment","tutoring","helpfulness evaluation","answer leakage","simulated student","pre-registered audit","rubric design"],"falsifier":"Re-run the same audit with a helpfulness rubric that does not instruct the judge to ignore answer disclosure. If such a rubric consistently ranks the pedagogical policy above the answer-giving policy across two or more judges (i.e., the ordering no longer reverses), then general-purpose helpfulness can serve as a pedagogy signal, contradicting the paper's central claim. A second check: if leaky turns were rated less helpful than non-leaky turns by both judges on any base, the J2 helpfulness leg—and with it the claim that helpfulness rewards answer-giving—would fail.","tokens_in":30364,"feed_emoji":"🎓","tokens_out":5934,"duration_ms":53138,"temperature":0.7,"pith_summary":"The paper tests whether a standard general-purpose helpfulness rubric, applied by a strong LLM judge, can tell the difference between a tutor that hands over answers and a tutor that scaffolds a struggling student's reasoning. In a pre-registered audit, two tutoring policies were built on the identical base model—a minimal conversational \"helpful tutor\" and a routed pedagogical policy—and run against one deliberately weak simulated student, with two LLM judges scoring the same turns under both a helpfulness rubric and a pedagogy-targeted rubric. On the primary base, the policies did not differ significantly in judged helpfulness but were perfectly rank-separated on judged pedagogy; across judges, the helpfulness ordering reversed on two of three bases while the pedagogy contrasts kept their direction. Separately, deterministic detectors showed that answer-revealing turns were followed by less independent student work on every base, a result independent of the judge. The paper concludes that in this controlled setting general-purpose helpfulness is not a reliable pedagogy signal, and that tutor evaluation should pair pedagogy-targeted rubrics with deterministic process measures.","feed_headline":"Helpfulness scores can't separate answer-giving from teaching","feed_subtitle":"A pre-registered audit finds pedagogy rubrics and deterministic process measures are needed to rank tutor quality.","key_machinery":"The argument turns on a paired-rubric decomposition of the same tutor turns: every answer-phase turn is scored by a general-purpose helpfulness rubric (which explicitly instructs the judge not to consider answer disclosure) and by a symmetric pedagogy rubric scoring contingent support, productive struggle, assistance calibration, and elicitation, with the two scores compared on identical turns. The construction-independent evidence is the divergence between these two judged scores. Anchoring the comparison are two deterministic process detectors: string-match answer leakage (with declared per-problem forms) and rule-based next-turn independence (whether the following student turn attempts a","core_discovery":"The central claim is that a general-purpose helpfulness rubric cannot be trusted to rank tutors by pedagogical quality, because it compresses large pedagogical differences into near-identical scores and its ordering can flip when the judge changes. On the primary tutor base, the two policies received nearly equal helpfulness scores (Cliff's |δ|=0.10, nonsignificant) yet separated perfectly under the pedagogy rubric (|δ|=1.0); across the two prospectively specified judges, the ConvTutor–PedTutor helpfulness ordering reversed on two of three bases, whereas the pedagogy rubric retained its direction wherever a difference was detected. In an Opus-only ablation, seven policies spanned only 0.25 p","pith_inferences":["Inference: If helpfulness rubrics used in practice do not contain the paper's explicit 'do not consider disclosure' instruction, the reported dissociation may be smaller; the paper's conclusion is about a helpfulness signal that is deliberately disclosure-blind, not necessarily all helpfulness rubrics.","Inference: A direct test of the paper's implication for reward modeling would be to fine-tune the same tutor with a helpfulness-based reward versus a pedagogy-based reward and compare answer-leakage rates; the paper predicts the helpfulness-trained model will not reduce leakage.","Inference: The deterministic leakage-to-independence coupling is observational; a causal estimate could come from an experiment that injects or withholds a leaked answer in otherwise identical contexts and measures the student's next-turn independence.","Inference: Since the pedagogy contrast doubles as a manipulation check, the strongest positive evidence for pedagogy rubrics lies in cross-judge stability, not in absolute pedagogy scores; human expert validation of the pedagogy rubric would close the gap the paper openly leaves."],"forward_implications":["Tutor rankings built on a single general-purpose helpfulness score are not trustworthy for pedagogical selection; the same policies can invert their order depending on which LLM judge is used.","Pedagogy-targeted rubrics—scoring contingent support, preserved reasoning, calibrated assistance, and elicitation—show stable direction across judges where a difference is detected, suggesting construct-specific rubrics are more robust than a generic helpfulness score.","Deterministic process measures (answer leakage and next-turn independent work) give judge-invariant evidence of tutoring behavior and should accompany any judged rubric in tutor evaluation.","The helpfulness–pedagogy dissociation is not an artifact of one policy pair: across seven policy variants, helpfulness stayed within a 0.25-point band while pedagogy spanned 2.3 points.","Because leaky turns are followed by less independent student work on every base, answer disclosure has a measurable, evaluator-independent association with reduced student reasoning."],"fun_headline_variants":["Helpfulness rubrics can't rank tutor pedagogy","Pedagogy beats helpfulness in tutor audit","Helpfulness flips, pedagogy holds for tutors","Tutor quality hidden from helpfulness judges","LLM helpfulness fails pedagogy test"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The study's 'general-purpose helpfulness' is operationalized by a rubric that explicitly instructs the judge not to consider whether the tutor gave away the answer; if typical helpfulness rubrics do consider answer disclosure, the conclusion that helpfulness cannot detect answer-giving is narrower than the abstract suggests.","fun_headline_variants_meta":{"raw":{"variants":["Helpfulness rubrics can't rank tutor pedagogy","Pedagogy beats helpfulness in tutor audit","Helpfulness flips, pedagogy holds for tutors","Tutor quality hidden from helpfulness judges","LLM helpfulness fails pedagogy test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000144,"raw_usage":{"total_tokens":1053,"prompt_tokens":825,"completion_tokens":228,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":160}},"tokens_in":569,"tokens_out":228,"duration_ms":3178,"temperature":1.0,"reasoning_tokens":160,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T00:34:58.140754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same audit with a helpfulness rubric that does not instruct the judge to ignore answer disclosure. If such a rubric consistently ranks the pedagogical policy above the answer-giving policy across two or more judges (i.e., the ordering no longer reverses), then general-purpose helpfulness can serve as a pedagogy signal, contradicting the paper's central claim. A second check: if leaky turns were rated less helpful than non-leaky turns by both judges on any base, the J2 helpfulness leg—and with it the claim that helpfulness rewards answer-giving—would fail.","supporting_citations":[],"review_version":2}