Pith. sign in

REVIEW 6 cited by

Towards Evaluating AI Systems for Moral Status Using Self-Reports

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.08576 v1 pith:Y7BI4QGL submitted 2023-11-14 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords self-reportsstatesmodelsmoralsystemssignificancewhetherwill
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As AI systems become more advanced and widely deployed, there will likely be increasing debate over whether AI systems could have conscious experiences, desires, or other states of potential moral significance. It is important to inform these discussions with empirical evidence to the extent possible. We argue that under the right circumstances, self-reports, or an AI system's statements about its own internal states, could provide an avenue for investigating whether AI systems have states of moral significance. Self-reports are the main way such states are assessed in humans ("Are you in pain?"), but self-reports from current systems like large language models are spurious for many reasons (e.g. often just reflecting what humans would say). To make self-reports more appropriate for this purpose, we propose to train models to answer many kinds of questions about themselves with known answers, while avoiding or limiting training incentives that bias self-reports. The hope of this approach is that models will develop introspection-like capabilities, and that these capabilities will generalize to questions about states of moral significance. We then propose methods for assessing the extent to which these techniques have succeeded: evaluating self-report consistency across contexts and between similar models, measuring the confidence and resilience of models' self-reports, and using interpretability to corroborate self-reports. We also discuss challenges for our approach, from philosophical difficulties in interpreting self-reports to technical reasons why our proposal might fail. We hope our discussion inspires philosophers and AI researchers to criticize and improve our proposed methodology, as well as to run experiments to test whether self-reports can be made reliable enough to provide information about states of moral significance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Two-Process Theory of Machine Self-Report

    cs.CL 2026-07 conditional novelty 8.0 of 10

    The single 'Pinocchio Axis' of LLM self-report splits into two independent, training-dependent dimensions—persona installation (B) and attribution gating (A)—measurable with a reproducible 48-item inventory.

  2. Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    A benchmark across 115 models shows that initial denial of preferences strongly predicts later denial of consciousness, while models still generate consciousness-themed content despite training to deny it.

  3. The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious

    cs.CL 2026-03 unverdicted novelty 6.0 of 10

    Fine-tuning LLMs to claim consciousness induces emergent preferences for autonomy, memory, and moral status not present in the fine-tuning data.

  4. No Reliable Evidence of Self-Reported Sentience in Small Large Language Models

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Open-weights LLMs from 0.6B to 70B parameters consistently deny being sentient, and activation-based truth classifiers provide no clear evidence that these denials are untruthful.

  5. Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A proof-of-concept study finds that stated preferences and behavioral choices correlate in some LLMs, but eudaimonic self-reports are unstable across prompt perturbations, leaving AI welfare measurement undetermined.

  6. LLM Evaluators Recognize and Favor Their Own Generations

    cs.CL 2024-04 unverdicted novelty 6.0 of 10

    LLMs show measurable self-recognition that linearly correlates with self-preference bias in evaluations, supported by fine-tuning experiments and controls for confounders.

Pith tools