Pith. sign in

REVIEW 7 cited by

GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10104 v2 pith:I6PDXMAK submitted 2024-02-15 cs.AI cs.CL

classification cs.AIcs.CL
keywords problemssubsetmodelsllmsbenchmarkgeometryaccuracybeen
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advancements in large language models (LLMs) and multi-modal models (MMs) have demonstrated their remarkable capabilities in problem-solving. Yet, their proficiency in tackling geometry math problems, which necessitates an integrated understanding of both textual and visual information, has not been thoroughly evaluated. To address this gap, we introduce the GeoEval benchmark, a comprehensive collection that includes a main subset of 2,000 problems, a 750 problems subset focusing on backward reasoning, an augmented subset of 2,000 problems, and a hard subset of 300 problems. This benchmark facilitates a deeper investigation into the performance of LLMs and MMs in solving geometry math problems. Our evaluation of ten LLMs and MMs across these varied subsets reveals that the WizardMath model excels, achieving a 55.67\% accuracy rate on the main subset but only a 6.00\% accuracy on the hard subset. This highlights the critical need for testing models against datasets on which they have not been pre-trained. Additionally, our findings indicate that GPT-series models perform more effectively on problems they have rephrased, suggesting a promising method for enhancing model capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective

    cs.AI 2025-06 conditional novelty 6.0 of 10

    When math problems are probed through nine human-cognitive dimensions, seven leading LLMs' effective pass rates drop 30 to 40 points below standard accuracy.

  2. Context-DPO: Aligning Language Models for Context-Faithfulness

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Context-DPO fine-tunes LLMs with direct preference optimization on counterfactual passages, yielding 35-280% context-faithfulness gains on its new ConFiQA benchmark.

  3. MageBench: Bridging Large Multimodal Models to Agents

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MageBench introduces a 483-scenario benchmark showing current large multimodal models are far weaker than humans at agent tasks requiring continuous visual feedback and planning.

  4. Towards Geometry Problem Solving in the Large Model Era: A Survey

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey that organizes geometry problem-solving research into benchmark construction, parsing, and reasoning, and proposes a unified parse-then-reason paradigm for the large-model era.

  5. Large Language Models for Mathematical Analysis

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A dataset of proof-based real analysis problems plus a classification, retrieval, and fine-tuning framework improves small LLMs' scores on that dataset, as judged by GPT-4o.

  6. Visual Large Language Models for Generalized and Specialized Applications

    cs.CV 2025-01 conditional novelty 3.0 of 10

    This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.

  7. Large Language Models as Computable Approximations to Solomonoff Induction

    cs.LG 2025-05 reject novelty 2.0 of 10

    The paper argues LLMs are computable approximations of Solomonoff induction, but its central derivation recovers the model's own probabilities by construction.

Pith tools