Pith. sign in

REVIEW 2 cited by

Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.18070 v2 pith:EC32G56I submitted 2024-01-31 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords llmsbiasescognitivemodelssolvingstepsarithmeticchildren
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There is increasing interest in employing large language models (LLMs) as cognitive models. For such purposes, it is central to understand which properties of human cognition are well-modeled by LLMs, and which are not. In this work, we study the biases of LLMs in relation to those known in children when solving arithmetic word problems. Surveying the learning science literature, we posit that the problem-solving process can be split into three distinct steps: text comprehension, solution planning and solution execution. We construct tests for each one in order to understand whether current LLMs display the same cognitive biases as children in these steps. We generate a novel set of word problems for each of these tests, using a neuro-symbolic approach that enables fine-grained control over the problem features. We find evidence that LLMs, with and without instruction-tuning, exhibit human-like biases in both the text-comprehension and the solution-planning steps of the solving process, but not in the final step, in which the arithmetic expressions are executed to obtain the answer.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

    cs.LG 2025-07 reject novelty 6.0 of 10

    A dynamic red-teaming audit reports that 94% of MedQA-correct answers fail under adversarial mutation, with 86% privacy leak rates, 81% bias shift rates, and 66-74% hallucination rates across 15 medical LLMs.

  2. Reasoning Strategies in Large Language Models: Can They Follow, Prefer, and Optimize?

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Prompting LLMs with distinct reasoning strategies and ensembling their outputs improves accuracy on logical deduction tasks, though not as consistently as the paper claims.

Pith tools