Pith. sign in

REVIEW 4 cited by

Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10648 v3 pith:XXDXCNUP submitted 2023-10-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords expertdecisionsbridgedatasetdecision-makingllmsmistakesnovice-expert
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scaling high-quality tutoring remains a major challenge in education. Due to growing demand, many platforms employ novice tutors who, unlike experienced educators, struggle to address student mistakes and thus fail to seize prime learning opportunities. Our work explores the potential of large language models (LLMs) to close the novice-expert knowledge gap in remediating math mistakes. We contribute Bridge, a method that uses cognitive task analysis to translate an expert's latent thought process into a decision-making model for remediation. This involves an expert identifying (A) the student's error, (B) a remediation strategy, and (C) their intention before generating a response. We construct a dataset of 700 real tutoring conversations, annotated by experts with their decisions. We evaluate state-of-the-art LLMs on our dataset and find that the expert's decision-making model is critical for LLMs to close the gap: responses from GPT4 with expert decisions (e.g., "simplify the problem") are +76% more preferred than without. Additionally, context-sensitive decisions are critical to closing pedagogical gaps: random decisions decrease GPT4's response quality by -97% than expert decisions. Our work shows the potential of embedding expert thought processes in LLM generations to enhance their capability to bridge novice-expert knowledge gaps. Our dataset and code can be found at: \url{https://github.com/rosewang2008/bridge}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration

    cs.AI 2025-06 conditional novelty 7.0 of 10

    Model benchmark performance only weakly predicts how well people learn from AI explanations, with notable outliers across code and math.

  2. IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations

    cs.CL 2025-09 conditional novelty 5.0 of 10

    IDEAlign uses a pick-the-odd-one-out triplet task to measure idea-level similarity between LLMs and expert human annotations, and shows LLM judges using this protocol outperform lexical and vector-based baselines.

  3. Intent Matters: Enhancing AI Tutoring with Fine-Grained Pedagogical Intent Annotation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Fine-tuning a math tutor model on 11 fine-grained pedagogical intents instead of 4 broad ones gave better automatic scores and a modest human preference in a small evaluation.

  4. SingaKids: A Multilingual Multimodal Dialogic Tutor for Language Learning

    cs.CL 2025-06 conditional novelty 4.0 of 10

    The paper presents SingaKids, a four-language dialogic tutoring system, and reports component-level improvements plus a 35-student pilot study of its scaffolding behavior.

Pith tools