Pith. sign in

REVIEW 26 cited by

Large Language Models for Education: A Survey and Outlook

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.18105 v2 pith:4DAMCQEI submitted 2024-03-26 cs.CL cs.AI

Large Language Models for Education: A Survey and Outlook

classification cs.CL cs.AI
keywords llmseducationsurveyeducationallanguagelargelearningmodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The advent of Large Language Models (LLMs) has brought in a new era of possibilities in the realm of education. This survey paper summarizes the various technologies of LLMs in educational settings from multifaceted perspectives, encompassing student and teacher assistance, adaptive learning, and commercial tools. We systematically review the technological advancements in each perspective, organize related datasets and benchmarks, and identify the risks and challenges associated with deploying LLMs in education. Furthermore, we outline future research opportunities, highlighting the potential promising directions. Our survey aims to provide a comprehensive technological picture for educators, researchers, and policymakers to harness the power of LLMs to revolutionize educational practices and foster a more effective personalized learning environment.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

    cs.CL 2024-10 unverdicted novelty 8.0

    ErrorRadar is a new benchmark of 2,500 multimodal K-12 math problems for MLLM error step identification and categorization, where GPT-4o trails human experts by ~10%.

  2. AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher Education

    cs.HC 2026-06 accept novelty 7.0

    Randomized experiment finds AI draft assistance raises feedback provision by teaching assistants 10.8 percentage points without harming quality.

  3. Towards Personalizing Secure Programming Education with LLM-Injected Vulnerabilities

    cs.CR 2026-04 conditional novelty 7.0

    LLM agents inject CWEs into student-authored code to generate personalized security examples; in a 71-student deployment, participants rated them more relevant than textbook cases but quantitative differences remained...

  4. Designing Safe and Accountable GenAI as a Learning Companion with Women Banned from Formal Education

    cs.CY 2026-04 conditional novelty 7.0

    Participatory design with 20 Afghan women reveals that safe GenAI learning companions must prioritize privacy, cultural fit, and genuine learning support, with the process itself linked to higher aspirations and agency.

  5. MisEdu-RAG: A Misconception-Aware Dual-Hypergraph RAG for Novice Math Teachers

    cs.IR 2026-04 unverdicted novelty 7.0

    MisEdu-RAG builds concept and instance hypergraphs for two-stage retrieval of pedagogical knowledge and student errors, improving feedback quality on the MisstepMath benchmark by 10.95% token-F1 and up to 15.3% on res...

  6. EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers

    cs.AI 2026-08 conditional novelty 6.0

    EduZone is a new evaluation framework and 5.2K-prompt dataset showing that LLMs are substantially more vulnerable to education-specific risks and adaptive multi-turn attacks than to conventional safety risks.

  7. Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification

    cs.AI 2026-07 conditional novelty 6.0

    LGU models implication and incompatibility among LLM answers and reports consistent AUROC/AUARC gains over semantic entropy on QA benchmarks.

  8. Memory in the LLM Era: Modular Architectures and Strategies in a Unified Framework

    cs.CL 2026-04 unverdicted novelty 6.0

    A unified framework for LLM agent memory is benchmarked, with a new hybrid method outperforming state-of-the-art on standard tasks.

  9. Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution

    cs.AI 2025-12 unverdicted novelty 6.0

    ReMe enables LLM agents to evolve via multi-faceted experience distillation, context-adaptive reuse, and utility-based memory refinement, achieving new SOTA results on BFCL-V3 and AppWorld while letting an 8B model be...

  10. Limitations on Accurate, Trusted, Human-level Reasoning

    cs.LG 2025-09 unverdicted novelty 6.0

    An accurate and trusted AI system cannot achieve human-level reasoning because there exist tasks easily solvable by humans but not by the system.

  11. In-depth Analysis of Graph-based RAG in a Unified Framework

    cs.IR 2025-03 unverdicted novelty 6.0

    A unified framework and large-scale comparison of graph-based RAG methods on QA tasks yields new high-performing variants obtained by recombining existing components.

  12. ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented Generation

    cs.IR 2025-02 unverdicted novelty 6.0

    ArchRAG proposes attributed-community hierarchical indexing and LLM clustering to improve accuracy and lower token usage in graph-based retrieval-augmented generation.

  13. How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions

    cs.CR 2026-07 conditional novelty 5.0

    Attacks that break LLMs best are not the ones that improve safety most; a Shapley- and greedy-based framework that selects attack subsets by downstream defender utility outperforms attacker-centric and attribution-onl...

  14. Lect\=uraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching

    cs.CL 2026-06 unverdicted novelty 5.0

    LectūraAgents proposes a hierarchical multi-agent system with adaptive embodied teaching and the TASA algorithm for personalized AI-assisted learning, reporting gains in content quality, teaching actions, and personal...

  15. Tackling the Root of Misinformation by Teaching Laypeople about Logical Fallacies via Socratic Questioning and Critical Argumentation

    cs.AI 2026-05 unverdicted novelty 5.0

    LFTutor combines LLMs with intent-driven Socratic questioning to tutor laypeople on logical fallacies and outperforms standard LLM baselines in automatic and human evaluations.

  16. Double-Edged Sword or Sharp Tool? Designing and Evaluating Triadic LLM-Teacher Collaboration for K-12 Writing at Scale

    cs.AI 2026-05 unverdicted novelty 5.0

    A two-year deployment across 120 schools shows that LLM-teacher collaboration improves K-12 writing quality via labor division, with a ceiling effect from excessive LLM linguistic expansion.

  17. ARIA: Adaptive Retrieval Intelligence Assistant -- A Multimodal RAG Framework for Domain-Specific Engineering Education

    cs.IR 2026-02 conditional novelty 5.0

    ARIA is a multimodal RAG framework that filters domain-specific questions with 97.5% accuracy and outperforms ChatGPT-5 on pedagogical quality for a university civil engineering course.

  18. Enhancing Large Language Model-Based Systems for End-to-End Circuit Analysis Problem Solving

    cs.CY 2025-12 conditional novelty 5.0

    Hybrid pipeline using YOLO vision and ngspice verification raises circuit analysis accuracy from Gemini's 79.52% baseline to 97.59%, with similar gains on hand-drawn diagrams.

  19. Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective

    cs.AI 2025-11 conditional novelty 5.0

    The paper analyzes CPU bottlenecks in agentic AI serving, selects representative workloads, and demonstrates that CPU-aware scheduling optimizations COMB and MAS can reduce P50 latency by up to 1.7x and total latency ...

  20. Teaching Astronomy with Large Language Models

    physics.ed-ph 2025-06 unverdicted novelty 5.0

    Structured integration of LLMs in astronomy education, including a domain-specific tutor and documentation requirements, leads to improved AI literacy and reduced student reliance on AI over the semester.

  21. ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education

    cs.CY 2026-06 unverdicted novelty 4.0

    ELEVATE is a framework and prototype for deploying LLM-powered 3D avatar tutors locally on consumer hardware with a three-stratum design separating interaction, execution, and governance layers.

  22. Design Principles and Observable Indicators for AI-Enabled Pedagogical Accompaniment: Evidence from the Amico Dual-Mode Prototype in Italy and China

    cs.HC 2026-05 unverdicted novelty 4.0

    The Amico dual-mode prototype in Italy and China provides initial evidence for a set of design principles and observable indicators that enable AI to serve as a relational bridge supporting human-centered pedagogical ...

  23. Are LLMs Ready for Computer Science Education? A Cross-Domain, Cross-Lingual and Cognitive-Level Evaluation Using Professional Certification Exams

    cs.CY 2026-04 unverdicted novelty 4.0

    GPT-5 leads English CS certifications, Qwen-Plus leads Chinese ones, DeepSeek-R1 is most balanced, and Llama-3.3 lags in higher reasoning and robustness, with performance dropping on complex questions.

  24. Beyond Grading Accuracy: Exploring Alignment of TAs and LLMs

    cs.CY 2026-03 conditional novelty 4.0

    Six open-source LLMs reach up to 88.56% per-criterion accuracy and Pearson r≈0.80 versus TA grades on 92 UML class diagrams, supporting mixed-initiative grading.

  25. What is (H)CI: Why Does the "Human'' Matter?

    cs.HC 2026-04 unverdicted novelty 2.0

    A workshop proposal to reflect on HCI's core identity and the importance of human elements in the era of generative AI.

  26. Beyond Answers: How LLMs Can Pursue Strategic Thinking in Education

    cs.CY 2025-04 unverdicted novelty 2.0

    LLMs can improve education by serving as tutors and collaborators that help students develop resolving strategies instead of providing direct solutions.