Pith. sign in

REVIEW 6 cited by

Retrieval-augmented Generation to Improve Math Question-Answering: Trade-offs Between Groundedness and Human Preference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.03184 v2 pith:RRU5E6UH submitted 2023-10-04 cs.CL cs.HC

classification cs.CLcs.HC
keywords responsesmathcontenteducationalgenerationimproveinteractivemiddle-school
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

For middle-school math students, interactive question-answering (QA) with tutors is an effective way to learn. The flexibility and emergent capabilities of generative large language models (LLMs) has led to a surge of interest in automating portions of the tutoring process - including interactive QA to support conceptual discussion of mathematical concepts. However, LLM responses to math questions can be incorrect or mismatched to the educational context - such as being misaligned with a school's curriculum. One potential solution is retrieval-augmented generation (RAG), which involves incorporating a vetted external knowledge source in the LLM prompt to increase response quality. In this paper, we designed prompts that retrieve and use content from a high-quality open-source math textbook to generate responses to real student questions. We evaluate the efficacy of this RAG system for middle-school algebra and geometry QA by administering a multi-condition survey, finding that humans prefer responses generated using RAG, but not when responses are too grounded in the textbook content. We argue that while RAG is able to improve response quality, designers of math QA systems must consider trade-offs between generating responses preferred by students and responses closely matched to specific educational resources.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Sparse-autoencoder features from LLMs trigger automatic prompt reformulation, yielding consistent gains on mathematical reasoning and metaphor detection.

  2. The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States

    cs.CL 2024-12 conditional novelty 6.0 of 10

    HalluRAG provides a recency-controlled dataset for closed-domain hallucination detection and shows that intermediate activation values carry hallucination signals as strongly as contextualized embeddings.

  3. Aligning LLMs for the Classroom with Knowledge-Based Retrieval -- A Comparative RAG Study

    cs.AI 2025-09 conditional novelty 5.0 of 10

    In classroom question-answering, vector RAG (OpenAI) excels at fact lookup, GraphRAG Global at thematic questions, and GraphRAG Local at dense altered textbooks; a simple query router combines their strengths.

  4. A Retrieval-Augmented Generation Framework for Academic Literature Navigation in Data Science

    cs.IR 2024-12 conditional novelty 4.0 of 10

    A five-stage enhanced RAG pipeline for data science literature is reported to improve LLM-judged context relevance, though the evaluation is self-contained and not externally validated.

  5. Zero-Indexing Internet Search Augmented Generation for Large Language Models

    cs.IR 2024-11 conditional novelty 4.0 of 10

    An internet search augmented generation system with a trained parser LLM, mixed ranking, and an extractor LLM reports better answers and 21-47% lower generative input-token cost than two RAG baselines.

  6. A Survey on Large Language Models for Mathematical Reasoning

    cs.AI 2025-06 conditional novelty 1.0 of 10

    Recent advances in LLM mathematical reasoning are organized into comprehension and generation phases, covering methods from prompting to test-time scaling.

Pith tools