Pith. sign in

REVIEW 1 cited by

Assessing the Impact of Prompting Methods on ChatGPT's Mathematical Capabilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.15006 v2 pith:LPP2JO63 submitted 2023-12-22 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords promptingmathematicalmethodsanalysisenhancingchatgpt-3effectivenessllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study critically evaluates the efficacy of prompting methods in enhancing the mathematical reasoning capability of large language models (LLMs). The investigation uses three prescriptive prompting methods - simple, persona, and conversational prompting - known for their effectiveness in enhancing the linguistic tasks of LLMs. We conduct this analysis on OpenAI's LLM chatbot, ChatGPT-3.5, on extensive problem sets from the MATH, GSM8K, and MMLU datasets, encompassing a broad spectrum of mathematical challenges. A grading script adapted to each dataset is used to determine the effectiveness of these prompting interventions in enhancing the model's mathematical analysis power. Contrary to expectations, our empirical analysis reveals that none of the investigated methods consistently improves over ChatGPT-3.5's baseline performance, with some causing significant degradation. Our findings suggest that prompting strategies do not necessarily generalize to new domains, in this study failing to enhance mathematical performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering

    cs.SE 2024-12 conditional novelty 4.0 of 10

    A systematic mapping study of 20 papers shows LLMs are mainly used for coding and thematic analysis, with efficiency benefits but reliability, nuance, and privacy limitations.

Pith tools