Pith. sign in

REVIEW 9 cited by

Language models show human-like content effects on reasoning tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.07051 v4 pith:ZNQWSQ5M submitted 2022-07-14 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagereasoningmodelscontenthumanhumanslogicaltasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Reasoning is a key ability for an intelligent system. Large language models (LMs) achieve above-chance performance on abstract reasoning tasks, but exhibit many imperfections. However, human abstract reasoning is also imperfect. For example, human reasoning is affected by our real-world knowledge and beliefs, and shows notable "content effects"; humans reason more reliably when the semantic content of a problem supports the correct logical inferences. These content-entangled reasoning patterns play a central role in debates about the fundamental nature of human intelligence. Here, we investigate whether language models $\unicode{x2014}$ whose prior expectations capture some aspects of human knowledge $\unicode{x2014}$ similarly mix content into their answers to logical problems. We explored this question across three logical reasoning tasks: natural language inference, judging the logical validity of syllogisms, and the Wason selection task. We evaluate state of the art large language models, as well as humans, and find that the language models reflect many of the same patterns observed in humans across these tasks $\unicode{x2014}$ like humans, models answer more accurately when the semantic content of a task supports the logical inferences. These parallels are reflected both in answer patterns, and in lower-level features like the relationship between model answer distributions and human response times. Our findings have implications for understanding both these cognitive effects in humans, and the factors that contribute to language model performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 53 citations worldwide. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs

    cs.CL 2025-07 conditional novelty 7.0 of 10

    Cognitive biases in LLMs are largely set during pretraining, while finetuning data and seed randomness only modulate them.

  3. Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Learned soft prefixes reliably flip correct syllogistic judgments in LLMs, transferring across unseen forms and interfaces and behaving mainly as a broad answer preference rather than a transferable logical operation.

  4. Enhancing Computational Cognitive Architectures with LLMs: A Case Study

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A case-study design shows how LLMs can serve as the implicit level of the Clarion cognitive architecture, giving it natural language and broader knowledge.

  5. Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Vision language models prompted with a low vision participant's vision profile and one example response reach only 70% agreement with that participant's held-out image answers.

  6. Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LLMs perform much worse on causal questions built from post-cutoff news articles, suggesting their apparent causal skill is mostly memorization, and a general-knowledge prompt method only partly closes the gap.

  7. MARS: Multi-hop Adaptive Retrieval and SPARQL Generation for KGQA

    cs.CL 2026-07 conditional novelty 5.0 of 10

    MARS answers multi-hop knowledge-graph questions by iteratively retrieving ranked triple patterns and letting an LLM decide when to emit a SPARQL query, beating agentic baselines on QALD-10 without fine-tuning.

  8. Towards General Continuous Memory for Vision-Language Models

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A vision-language model can act as its own continuous memory encoder, compressing external multimodal knowledge into eight embeddings that improve reasoning when prepended to the frozen model.

  9. Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A survey that classifies chain-of-thought methods for autonomous driving into modular, logical, and reflective pipelines, and proposes three evolutionary stages from direct prompting to reinforcement learning.

Pith tools