Open-weights LLMs from 0.6B to 70B parameters consistently deny being sentient, and activation-based truth classifiers provide no clear evidence that these denials are untruthful.
URL https: //link.springer.com/10.1007/s11098-025-02343-7
1 Pith paper cite this work, alongside 5 external citations. Polarity classification is still indexing.
1
Pith paper citing it
5
external citations · OpenAlex
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
No Reliable Evidence of Self-Reported Sentience in Small Large Language Models
Open-weights LLMs from 0.6B to 70B parameters consistently deny being sentient, and activation-based truth classifiers provide no clear evidence that these denials are untruthful.