Pith. sign in

REVIEW 3 cited by

A Systematic Review on Prompt Engineering in Large Language Models for K-12 STEM Education

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.11123 v1 pith:RIVJLESK submitted 2024-10-14 cs.CL cs.HC

A Systematic Review on Prompt Engineering in Large Language Models for K-12 STEM Education

classification cs.CL cs.HC
keywords modelsllmspromptpromptingeducationengineeringk-12stem
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models (LLMs) have the potential to enhance K-12 STEM education by improving both teaching and learning processes. While previous studies have shown promising results, there is still a lack of comprehensive understanding regarding how LLMs are effectively applied, specifically through prompt engineering-the process of designing prompts to generate desired outputs. To address this gap, our study investigates empirical research published between 2021 and 2024 that explores the use of LLMs combined with prompt engineering in K-12 STEM education. Following the PRISMA protocol, we screened 2,654 papers and selected 30 studies for analysis. Our review identifies the prompting strategies employed, the types of LLMs used, methods of evaluating effectiveness, and limitations in prior work. Results indicate that while simple and zero-shot prompting are commonly used, more advanced techniques like few-shot and chain-of-thought prompting have demonstrated positive outcomes for various educational tasks. GPT-series models are predominantly used, but smaller and fine-tuned models (e.g., Blender 7B) paired with effective prompt engineering outperform prompting larger models (e.g., GPT-3) in specific contexts. Evaluation methods vary significantly, with limited empirical validation in real-world settings.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Practice Less, Explain More: LLM-Supported Self-Explanation Improves Explanation Quality on Transfer Problems in Calculus

    cs.HC 2026-03 conditional novelty 5.0

    LLM-supported open-ended self-explanation improved explanation quality on Not-Enough-Information calculus transfer problems versus control, with no post-test accuracy gains and far fewer practice problems completed.

  2. Teaching Astronomy with Large Language Models

    physics.ed-ph 2025-06 unverdicted novelty 5.0

    Structured integration of LLMs in astronomy education, including a domain-specific tutor and documentation requirements, leads to improved AI literacy and reduced student reliance on AI over the semester.

  3. Comparing RAG and GraphRAG for Page-Level Retrieval Question Answering on a Math Textbook

    cs.IR 2025-09 conditional novelty 4.0

    On a 477-question page-level math textbook benchmark, embedding-based RAG with voyage-3-large reaches 99.4% top-10 retrieval accuracy and outperforms GraphRAG for retrieval and answer quality.