Pith. sign in

REVIEW 6 cited by

Language Models as Science Tutors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.11111 v2 pith:24YK2SV2 submitted 2024-02-16 cs.CL

classification cs.CL
keywords modelstutorevallongscientifictutorchatbenchmarkdatasetslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

NLP has recently made exciting progress toward training language models (LMs) with strong scientific problem-solving skills. However, model development has not focused on real-life use-cases of LMs for science, including applications in education that require processing long scientific documents. To address this, we introduce TutorEval and TutorChat. TutorEval is a diverse question-answering benchmark consisting of questions about long chapters from STEM textbooks, written by experts. TutorEval helps measure real-life usability of LMs as scientific assistants, and it is the first benchmark combining long contexts, free-form generation, and multi-disciplinary scientific knowledge. Moreover, we show that fine-tuning base models with existing dialogue datasets leads to poor performance on TutorEval. Therefore, we create TutorChat, a dataset of 80,000 long synthetic dialogues about textbooks. We use TutorChat to fine-tune Llemma models with 7B and 34B parameters. These LM tutors specialized in math have a 32K-token context window, and they excel at TutorEval while performing strongly on GSM8K and MATH. Our datasets build on open-source materials, and we release our models, data, and evaluations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification

    cs.RO 2026-03 conditional novelty 7.0 of 10

    LLMs struggle on research-level robotic path-planning approximation proofs unless given task-specific lemmas, which help more than CoT or oracle ratios.

  2. When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration

    cs.AI 2025-06 conditional novelty 7.0 of 10

    Model benchmark performance only weakly predicts how well people learn from AI explanations, with notable outliers across code and math.

  3. Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment

    cs.AI 2026-06 reject novelty 6.0 of 10

    SPRA reduces scaffolding collapse in LLM Socratic tutors from ~68% to 32% collapse rate by adding a margin-preserving representation loss aligned to frozen reference hidden states.

  4. LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents

    cs.CY 2026-06 conditional novelty 6.0 of 10

    A benchmark dataset pairing 647 human-written STEM lessons with generated, human-reviewed lesson plans across 240 topics, plus a proposed three-part evaluation pipeline.

  5. ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations

    cs.CL 2025-06 conditional novelty 6.0 of 10

    ELI-Why shows GPT-4's grade-tailored explanations often miss the intended educational level and are less informative than human-curated explanations.

  6. Robust pid sliding mode control for dc servo motor speed control

    eess.SY 2025-08 unverdicted novelty 2.0 of 10

    An abstract-only claim that SMC-PID outperforms PID for DC servo motor speed on the CE110 trainer; the submitted body text is an unrelated paper, so the result is unverifiable.

Pith tools