Pith. sign in

REVIEW 8 cited by

How well do Large Language Models perform in Arithmetic tasks?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.02015 v1 pith:J2ZGI2I7 submitted 2023-03-16 cs.CL cs.AI

How well do Large Language Models perform in Arithmetic tasks?

classification cs.CL cs.AI
keywords arithmeticlanguagelargemodelsmathproblemsstepabilities
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models have emerged abilities including chain-of-thought to answer math word problems step by step. Solving math word problems not only requires abilities to disassemble problems via chain-of-thought but also needs to calculate arithmetic expressions correctly for each step. To the best of our knowledge, there is no work to focus on evaluating the arithmetic ability of large language models. In this work, we propose an arithmetic dataset MATH 401 to test the latest large language models including GPT-4, ChatGPT, InstrctGPT, Galactica, and LLaMA with various arithmetic expressions and provide a detailed analysis of the ability of large language models. MATH 401 and evaluation codes are released at \url{https://github.com/GanjinZero/math401-llm}.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models

    cs.LG 2026-05 unverdicted novelty 7.0

    LLM residual streams during addition form an Iso-Raw-Sum Trajectory anchored by digit semantics and modulated by continuous carry signals, with errors arising as geometric slippages across quantization thresholds in a...

  2. DEL: Digit Entropy Loss for Numerical Learning of Large Language Models

    cs.CL 2026-05 conditional novelty 6.0

    DEL is a new loss for LLM numerical learning that applies supervised digit entropy optimization and extends to floating-point numbers, showing improved accuracy and distance metrics over prior methods on math benchmarks.

  3. Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs

    cs.CL 2026-04 unverdicted novelty 6.0

    Multimodal LLMs perceive numbers accurately across modalities but fail at multi-digit multiplication, with performance predicted by an arithmetic load metric C and degradation confirmed as computational rather than pe...

  4. TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

    cs.RO 2026-02 conditional novelty 6.0

    The log-probability a VLM assigns to 'True' for 'does this video prefix complete the task?' is used as a zero-shot dense progress reward that outperforms GVL on open-source models.

  5. Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

    cs.CL 2023-08 unverdicted novelty 6.0

    Pre-training loss predicts LLM math reasoning better than parameter count; rejection sampling fine-tuning with diverse paths raises LLaMA-7B accuracy on GSM8K from 35.9% with SFT to 49.3%.

  6. Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization

    cs.RO 2026-04 unverdicted novelty 5.0

    EEAgent with LSTRO sets new state-of-the-art results on six VIMA-Bench robotic manipulation tasks by dynamically refining prompts through reflection on successes and failures.

  7. Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

    cs.CL 2026-05 accept novelty 4.0

    A structured survey of LLM mathematical reasoning that unifies dataset taxonomies, reviews architectures and training strategies, and highlights the gap between answer accuracy and process-level verification.

  8. Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

    cs.CL 2026-05 unverdicted novelty 3.0

    A literature survey synthesizing benchmarks, architectures, training strategies, and evaluation methods for mathematical reasoning in LLMs, based on roughly 120 papers.