REVIEW 6 cited by
Retrieval-augmented Generation to Improve Math Question-Answering: Trade-offs Between Groundedness and Human Preference
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
For middle-school math students, interactive question-answering (QA) with tutors is an effective way to learn. The flexibility and emergent capabilities of generative large language models (LLMs) has led to a surge of interest in automating portions of the tutoring process - including interactive QA to support conceptual discussion of mathematical concepts. However, LLM responses to math questions can be incorrect or mismatched to the educational context - such as being misaligned with a school's curriculum. One potential solution is retrieval-augmented generation (RAG), which involves incorporating a vetted external knowledge source in the LLM prompt to increase response quality. In this paper, we designed prompts that retrieve and use content from a high-quality open-source math textbook to generate responses to real student questions. We evaluate the efficacy of this RAG system for middle-school algebra and geometry QA by administering a multi-condition survey, finding that humans prefer responses generated using RAG, but not when responses are too grounded in the textbook content. We argue that while RAG is able to improve response quality, designers of math QA systems must consider trade-offs between generating responses preferred by students and responses closely matched to specific educational resources.
Forward citations
Cited by 6 Pith papers
-
Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
Sparse-autoencoder features from LLMs trigger automatic prompt reformulation, yielding consistent gains on mathematical reasoning and metaphor detection.
-
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
HalluRAG provides a recency-controlled dataset for closed-domain hallucination detection and shows that intermediate activation values carry hallucination signals as strongly as contextualized embeddings.
-
Aligning LLMs for the Classroom with Knowledge-Based Retrieval -- A Comparative RAG Study
In classroom question-answering, vector RAG (OpenAI) excels at fact lookup, GraphRAG Global at thematic questions, and GraphRAG Local at dense altered textbooks; a simple query router combines their strengths.
-
A Retrieval-Augmented Generation Framework for Academic Literature Navigation in Data Science
A five-stage enhanced RAG pipeline for data science literature is reported to improve LLM-judged context relevance, though the evaluation is self-contained and not externally validated.
-
Zero-Indexing Internet Search Augmented Generation for Large Language Models
An internet search augmented generation system with a trained parser LLM, mixed ranking, and an extractor LLM reports better answers and 21-47% lower generative input-token cost than two RAG baselines.
-
A Survey on Large Language Models for Mathematical Reasoning
Recent advances in LLM mathematical reasoning are organized into comprehension and generation phases, covering methods from prompting to test-time scaling.
Discussion (0). Continue with ORCID to comment.