Pith. sign in

REVIEW 2 cited by

Evaluating Large Language Models on the GMAT: Implications for the Future of Business Education

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.02985 v1 pith:AGFUTZ7S submitted 2024-01-02 cs.CL cs.AI

Evaluating Large Language Models on the GMAT: Implications for the Future of Business Education

classification cs.CL cs.AI
keywords modelsgpt-4turbobusinesseducationllmsclaudeapplication
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The rapid evolution of artificial intelligence (AI), especially in the domain of Large Language Models (LLMs) and generative AI, has opened new avenues for application across various fields, yet its role in business education remains underexplored. This study introduces the first benchmark to assess the performance of seven major LLMs, OpenAI's models (GPT-3.5 Turbo, GPT-4, and GPT-4 Turbo), Google's models (PaLM 2, Gemini 1.0 Pro), and Anthropic's models (Claude 2 and Claude 2.1), on the GMAT, which is a key exam in the admission process for graduate business programs. Our analysis shows that most LLMs outperform human candidates, with GPT-4 Turbo not only outperforming the other models but also surpassing the average scores of graduate students at top business schools. Through a case study, this research examines GPT-4 Turbo's ability to explain answers, evaluate responses, identify errors, tailor instructions, and generate alternative scenarios. The latest LLM versions, GPT-4 Turbo, Claude 2.1, and Gemini 1.0 Pro, show marked improvements in reasoning tasks compared to their predecessors, underscoring their potential for complex problem-solving. While AI's promise in education, assessment, and tutoring is clear, challenges remain. Our study not only sheds light on LLMs' academic potential but also emphasizes the need for careful development and application of AI in education. As AI technology advances, it is imperative to establish frameworks and protocols for AI interaction, verify the accuracy of AI-generated content, ensure worldwide access for diverse learners, and create an educational environment where AI supports human expertise. This research sets the stage for further exploration into the responsible use of AI to enrich educational experiences and improve exam preparation and assessment methods.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns

    cs.SE 2026-06 unverdicted novelty 4.0

    Gemini 3 Flash achieved the highest accuracy on PSM I-style questions among three tested LLMs, with low intra-model variability and systematic error patterns by question format and topic.

  2. Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study

    cs.SE 2026-06 unverdicted novelty 3.0

    GPT-5 with source-citation prompting achieves 89.1% accuracy on 993 PSM questions, outperforming zero-shot and chain-of-thought while errors cluster in multi-select and interpretive topics.