Pith. sign in

Testing GPT-4 with Wolfram Alpha and Code Interpreter plug-ins on math and science problems

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

This report describes a test of the large language model GPT-4 with the Wolfram Alpha and the Code Interpreter plug-ins on 105 original problems in science and math, at the high school and college levels, carried out in June-August 2023. Our tests suggest that the plug-ins significantly enhance GPT's ability to solve these problems. Having said that, there are still often "interface" failures; that is, GPT often has trouble formulating problems in a way that elicits useful answers from the plug-ins. Fixing these interface failures seems like a central challenge in making GPT a reliable tool for college-level calculation problems.

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

unclear 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Evaluation of LLMs for mathematical problem solving cs.AI · 2025-05-30 · reject · none · ref 51 · internal anchor

    A three-model, three-dataset LLM math evaluation using a multi-dimensional reasoning rubric, undermined by contradictory accuracy tables.