Pith. sign in

REVIEW 10 cited by

Does It Make Sense to Speak of Introspection in Large Language Models?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.05068 v2 pith:HZ274JUW submitted 2025-06-05 cs.CL cs.AI

classification cs.CLcs.AI
keywords introspectionexamplellmsarguebehaviourlanguagelargelinguistic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) exhibit compelling linguistic behaviour, and sometimes offer self-reports, that is to say statements about their own nature, inner workings, or behaviour. In humans, such reports are often attributed to a faculty of introspection and are typically linked to consciousness. This raises the question of how to interpret self-reports produced by LLMs, given their increasing linguistic fluency and cognitive capabilities. To what extent (if any) can the concept of introspection be meaningfully applied to LLMs? Here, we present and critique two examples of apparent introspective self-report from LLMs. In the first example, an LLM attempts to describe the process behind its own "creative" writing, and we argue this is not a valid example of introspection. In the second example, an LLM correctly infers the value of its own temperature parameter, and we argue that this can be legitimately considered a minimal example of introspection, albeit one that is (presumably) not accompanied by conscious experience.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

    cs.AI 2026-07 conditional novelty 7.0 of 10

    LLMs' source-attribution ability is not fixed: it flips with conversational memory structure, and corrective feedback can invert judgments or sever confidence from accuracy.

  2. Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Strictly pre-answer hidden states of a looped transformer add significant AUROC over surface shortcuts for predicting correctness, and the readout yields decision-level gains but no generative control.

  3. Verbalizable Representations Form a Global Workspace in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Language models represent their current reasoning in a small, readable set of verbalizable vectors (the J-space) that functions like a global workspace.

  4. The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    The primary axis of psychometric variation among LLMs is the degree to which they represent themselves as loci of phenomenal experience rather than systems of behavioral responses.

  5. Asymmetric Communication: Large Language Models and Language Games

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Human–LLM exchange is asymmetric communication: model outputs circulate without commitments, so AGI, hallucination, agency, sentience, and alignment are receiver-side category mistakes, and alignment is institutional ...

  6. Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Fixed counterfactual explanation datasets train LMs such that generated explanations track the model's evolving behavior rather than the fixed targets, due to persistent correlation during training.

  7. Language models recognize dropout and Gaussian noise applied to their activations

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    Language models detect, localize, and distinguish dropout from Gaussian noise applied to their activations, often with high accuracy.

  8. Privileged Self-Access Matters for Introspection in AI

    cs.AI 2025-08 conditional novelty 6.0 of 10

    LLMs' temperature self-reports are confounded by prompt style and show no privileged self-access, supporting a thicker definition of AI introspection.

  9. On the Creativity of AI Agents

    cs.CY 2026-04 unverdicted novelty 5.0 of 10

    LLM agents produce outputs that meet basic functional criteria for creativity but lack the process-level, social, and personal elements required for ontological creativity.

  10. AI and Consciousness: Shifting Focus Towards Tractable Questions

    cs.CY 2026-05 unverdicted novelty 3.0 of 10

    Direct research on AI consciousness is intractable, so the field should prioritize studying perceived AI consciousness and its societal consequences.

Pith tools