Pith. sign in

REVIEW 21 cited by

Looking Inward: Language Models Can Learn About Themselves by Introspection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13787 v1 pith:7BYZMFDO submitted 2024-10-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelintrospectionbehaviorpredicteveninternalitselfmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind (e.g., thoughts and feelings) that is not accessible to external observers. Can LLMs introspect? We define introspection as acquiring knowledge that is not contained in or derived from training data but instead originates from internal states. Such a capability could enhance model interpretability. Instead of painstakingly analyzing a model's internal workings, we could simply ask the model about its beliefs, world models, and goals. More speculatively, an introspective model might self-report on whether it possesses certain internal states such as subjective feelings or desires and this could inform us about the moral status of these states. Such self-reports would not be entirely dictated by the model's training data. We study introspection by finetuning LLMs to predict properties of their own behavior in hypothetical scenarios. For example, "Given the input P, would your output favor the short- or long-term option?" If a model M1 can introspect, it should outperform a different model M2 in predicting M1's behavior even if M2 is trained on M1's ground-truth behavior. The idea is that M1 has privileged access to its own behavioral tendencies, and this enables it to predict itself better than M2 (even if M2 is generally stronger). In experiments with GPT-4, GPT-4o, and Llama-3 models (each finetuned to predict itself), we find that the model M1 outperforms M2 in predicting itself, providing evidence for introspection. Notably, M1 continues to predict its behavior accurately even after we intentionally modify its ground-truth behavior. However, while we successfully elicit introspection on simple tasks, we are unsuccessful on more complex tasks or those requiring out-of-distribution generalization.

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

    cs.AI 2026-07 conditional novelty 7.0 of 10

    LLMs' source-attribution ability is not fixed: it flips with conversational memory structure, and corrective feedback can invert judgments or sever confidence from accuracy.

  2. Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Strictly pre-answer hidden states of a looped transformer add significant AUROC over surface shortcuts for predicting correctness, and the readout yields decision-level gains but no generative control.

  3. Verbalizable Representations Form a Global Workspace in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Language models represent their current reasoning in a small, readable set of verbalizable vectors (the J-space) that functions like a global workspace.

  4. MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Non-invasive per-utterance belief probes in Mafia, auto-scored against engine truth, expose poorly calibrated LLM confidence and 1.5× over-prediction of being suspected.

  5. When Should We Protect AI? A Precautionary Framework for Consciousness Uncertainty

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    A precautionary framework with five consciousness dimensions, threshold-plus-gradation rules, and dual aggregation methods translates evidence into protective obligations for AI.

  6. ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    ContextEcho benchmark shows persona drift occurs across 23 frontier models in long agentic-coding sessions, is not reliably reset by compaction, and can be restored by single-shot anchors with mode-dependent effects.

  7. The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    The primary axis of psychometric variation among LLMs is the degree to which they represent themselves as loci of phenomenal experience rather than systems of behavioral responses.

  8. Asymmetric Communication: Large Language Models and Language Games

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Human–LLM exchange is asymmetric communication: model outputs circulate without commitments, so AGI, hallucination, agency, sentience, and alignment are receiver-side category mistakes, and alignment is institutional ...

  9. Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    SFT lessons — reason-based training, on-model replay, and wash-out robustness — transfer across toy models, model organisms, and alignment SFT, improving the capability–safety tradeoff.

  10. MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A released open-source testbed that privately probes LLM agents' beliefs during Mafia games shows agents are overconfident, over-predict suspicion by 1.5x, and rarely change outcomes when wrong beliefs lock in a vote.

  11. Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Experiments reveal that LLMs follow instructions at rates from 1% to 99% when opposed by hardcoded conflicting patterns, with robustness tied to output diversity and alignment with model priors rather than general capability.

  12. Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    LLMs show instruction-following rates from 1% to 99% when instructions conflict with hardcoded pattern demonstrations, with output diversity as the main predictor of resistance.

  13. Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Small LLMs can be fine-tuned to localize activation-steering perturbations, raising Llama-1B accuracy from 9.6% to 60.6% and generalizing to a strength-comparison task.

  14. Characterizing the Consistency of the Emergent Misalignment Persona

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    Fine-tuning LLMs on narrow misaligned data produces either coherent-persona models where harmful outputs match self-reported misalignment or inverted-persona models where harmful outputs occur alongside claims of alignment.

  15. Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    A benchmark across 115 models shows that initial denial of preferences strongly predicts later denial of consciousness, while models still generate consciousness-themed content despite training to deny it.

  16. No Reliable Evidence of Self-Reported Sentience in Small Large Language Models

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Open-weights LLMs from 0.6B to 70B parameters consistently deny being sentient, and activation-based truth classifiers provide no clear evidence that these denials are untruthful.

  17. The Assistant as a Privileged Persona: A canonical reference in cross-persona self-recognition

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    On Llama-3.1-70B-Instruct the Assistant persona functions as the sole canonical reference for cross-persona authorship judgments, with symmetric entropy gaps predicting only on its row and asymmetric surprise relative...

  18. Some[Body] Must Receive That Pain for Agent Accountability

    cs.CY 2026-05 unverdicted novelty 5.0 of 10

    AI agents lack the persistent identity and feedback mechanisms needed for consequence reception, requiring new architectures or continued human accountability.

  19. Phase Transitions in Driven Informational Systems: A Two-Field Perspective on Learning Theory and Non-Equilibrium Chemistry

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Proposes a two-gradient-field model with candidate order parameters alpha_dagger and kappa_c to unify phase transitions across learning theory and non-equilibrium chemistry.

  20. Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power

    cs.CY 2026-04 unverdicted novelty 5.0 of 10

    AI discourse employs strategically polysemous terms that blend technical precision with anthropomorphic implications, enabling glosslighting that sustains hype and deflects scrutiny.

  21. When Self-Reference Fails to Close: Matrix-Level Dynamics in Large Language Models

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    Non-closing truth recursion prompts destabilize LLM attention matrices with large effect sizes, unlike grounded self-reference or factual controls, and increase contradictory model outputs.

Pith tools