Gemini 3 Flash achieved the highest accuracy on PSM I-style questions among three tested LLMs, with low intra-model variability and systematic error patterns by question format and topic.
Evaluating large language models on the gmat: Implications for the future of business education,
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.SE 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
GPT-5 with source-citation prompting achieves 89.1% accuracy on 993 PSM questions, outperforming zero-shot and chain-of-thought while errors cluster in multi-select and interpretive topics.
citing papers explorer
-
Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns
Gemini 3 Flash achieved the highest accuracy on PSM I-style questions among three tested LLMs, with low intra-model variability and systematic error patterns by question format and topic.
-
Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study
GPT-5 with source-citation prompting achieves 89.1% accuracy on 993 PSM questions, outperforming zero-shot and chain-of-thought while errors cluster in multi-select and interpretive topics.