Pith. sign in

REVIEW 5 cited by

A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11050 v2 pith:GK6P2W4X submitted 2024-06-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords tokenbiasreasoningllmsgenuineabilitiesframeworkhypotheses
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study introduces a hypothesis-testing framework to assess whether large language models (LLMs) possess genuine reasoning abilities or primarily depend on token bias. We go beyond evaluating LLMs on accuracy; rather, we aim to investigate their token bias in solving logical reasoning tasks. Specifically, we develop carefully controlled synthetic datasets, featuring conjunction fallacy and syllogistic problems. Our framework outlines a list of hypotheses where token biases are readily identifiable, with all null hypotheses assuming genuine reasoning capabilities of LLMs. The findings in this study suggest, with statistical guarantee, that most LLMs still struggle with logical reasoning. While they may perform well on classic problems, their success largely depends on recognizing superficial patterns with strong token bias, thereby raising concerns about their actual reasoning and generalization abilities. Codes and data are open-sourced at https://github.com/bowen-upenn/llm_token_bias.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Implicit Reasoning Steering via Concept Chaining

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Reinforcement-learning-optimized concept-chain paragraphs covertly steer language-model multiple-choice preferences after continued pretraining, with far lower detectability than direct paraphrases.

  2. Understanding the Ability of LLMs to Handle Character-Level Perturbation

    cs.CL 2025-10 conditional novelty 6.0 of 10

    LLMs remain surprisingly accurate on math and coding when invisible Unicode noise is inserted after every character, with robustness driven by implicit internal denoising and, for some models, explicit rewriting in ch...

  3. Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence

    cs.AI 2025-06 conditional novelty 5.0 of 10

    Reliable deductive reasoning in AI requires replacing average-case statistical objectives with the exact learning criterion of universal correctness, a thesis supported by sample-complexity lower bounds showing statis...

  4. Do Large Language Models Reason Causally Like Us? Even Better?

    cs.AI 2025-02 conditional novelty 5.0 of 10

    GPT-4o, Gemini-Pro, and Claude show less associative bias than human reasoners on collider graphs, making them more normatively aligned in likelihood judgments, but none fully demonstrates explaining away.

  5. Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A symbol-remapping benchmark shows that LLMs' arithmetic and symbolic reasoning accuracy drops sharply when familiar digits and operators are replaced, revealing heavy reliance on memorized tokens rather than abstract rules.

Pith tools