REVIEW 6 cited by
Language Models as Science Tutors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
NLP has recently made exciting progress toward training language models (LMs) with strong scientific problem-solving skills. However, model development has not focused on real-life use-cases of LMs for science, including applications in education that require processing long scientific documents. To address this, we introduce TutorEval and TutorChat. TutorEval is a diverse question-answering benchmark consisting of questions about long chapters from STEM textbooks, written by experts. TutorEval helps measure real-life usability of LMs as scientific assistants, and it is the first benchmark combining long contexts, free-form generation, and multi-disciplinary scientific knowledge. Moreover, we show that fine-tuning base models with existing dialogue datasets leads to poor performance on TutorEval. Therefore, we create TutorChat, a dataset of 80,000 long synthetic dialogues about textbooks. We use TutorChat to fine-tune Llemma models with 7B and 34B parameters. These LM tutors specialized in math have a 32K-token context window, and they excel at TutorEval while performing strongly on GSM8K and MATH. Our datasets build on open-source materials, and we release our models, data, and evaluations.
Forward citations
Cited by 6 Pith papers
-
Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification
LLMs struggle on research-level robotic path-planning approximation proofs unless given task-specific lemmas, which help more than CoT or oracle ratios.
-
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
Model benchmark performance only weakly predicts how well people learn from AI explanations, with notable outliers across code and math.
-
Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment
SPRA reduces scaffolding collapse in LLM Socratic tutors from ~68% to 32% collapse rate by adding a margin-preserving representation loss aligned to frozen reference hidden states.
-
LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents
A benchmark dataset pairing 647 human-written STEM lessons with generated, human-reviewed lesson plans across 240 topics, plus a proposed three-part evaluation pipeline.
-
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
ELI-Why shows GPT-4's grade-tailored explanations often miss the intended educational level and are less informative than human-curated explanations.
-
Robust pid sliding mode control for dc servo motor speed control
An abstract-only claim that SMC-PID outperforms PID for DC servo motor speed on the CE110 trainer; the submitted body text is an unrelated paper, so the result is unverifiable.
Discussion (0). Sign in to comment.