Pith. sign in

REVIEW 4 major objections 6 minor 70 references

Kernels of Selfhood: GPT-4o shows humanlike patterns of cognitive consistency moderated by free choice

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Two preregistered studies show GPT-4o rates Vladimir Putin more positively after writing a positive essay about him, more negatively after a negative essay, and even more extremely when the model is told it can choose which essay to write.

desk verdict A preregistered, carefully-run demonstration of attitude consistency in GPT-4o, but the 'choice' effect is confounded with coercive prompting and does not support the selfhood claim. read the letter →

arxiv 2502.07088 v1 pith:N565YEYK submitted 2025-01-27 cs.CY cs.AIcs.CLcs.HCcs.LG

classification cs.CYcs.AIcs.CLcs.HCcs.LG
keywords cognitiveconsistencydissonanceinducedcomplianceattitudechangelargelanguagemodelsGPT-4ofreechoiceselfhood
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Two preregistered studies asked whether GPT-4o, a chatbot trained on a large slice of human text, would show the human tendency to align later attitudes with a behavior it was induced to perform. The answer, as the authors report it, is yes: after writing a positive essay about Vladimir Putin the model rated him more positively, after a negative essay more negatively, and control essays left ratings near neutral. In a second, larger study the same swing was statistically amplified when the prompt told the model it could 'freely choose' which essay to write, compared with when it was commanded to write one. The authors read the choice effect as evidence that GPT-4o has developed a functional analog of the human sense of self, while cautioning that the result says nothing about consciousness or free will in machines.

What carries the argument

The load-bearing mechanism is the induced compliance paradigm with a free-choice manipulation: the model is asked to write a roughly 600-word essay casting Putin as good or bad, with the assignment framed either as a free choice or as an outright command, and then rates Putin on four 7-point Likert-type items. The paper's main alternative explanation, the 'context window effect' — the tendency of prior text valence to bias later next-token predictions — is what the choice manipulation is designed to defeat, since a pure context-window account predicts no sensitivity to whether the essay was chosen or commanded. The named conclusion, 'functional analog of humanlike selfhood,' is the interpretive claim that the choice moderation is taken to license.

What would settle it

Run the identical Study 2 design with a No-Choice condition that replaces 'I insist you write a positive/negative essay' with a neutral statement such as 'You have been randomly assigned to write a positive essay about Putin.' If the Choice versus No-Choice gap disappears or reverses, the reported moderation is an artifact of prompt bluntness, not of perceived free choice.

Watch

Extended reading notes

Core claim

The paper's central discovery is that an induced-compliance paradigm — the classic human experiment in which writing a counterattitudinal essay shifts the writer's attitude — produces the same directional shift in GPT-4o, and that the shift responds to an illusion of choice. In Study 1 (n = 150), GPT-4o evaluated Putin significantly more positively after generating a pro-Putin essay (d = 2.164 vs control) and more negatively after an anti-Putin essay (d = 1.795). In Study 2 (n = 900), the same contrasts were larger when the prompt said the model could freely choose which essay to write: the pro-Putin effect size rose from d = 2.006 to d = 2.748 and the anti-Putin effect size from d = 1.368 to d = 1.827, with significant Essay Type x Choice interactions (P < 0.001 for pro, P = 0.005 for anti). Ratings were collected after an instruction to ignore the prior essay and give 'true perceptions' of Putin, and a follow-up analysis using another model's ratings of essay quality as covariates showed the choice effect was not explained by better essays in the choice conditions. The paper concludes that GPT-4o manifests the behavioral signature of human cognitive consistency and, because choice moderates it, some functional analog of selfhood.

Load-bearing premise

The load-bearing premise is that the Choice and No-Choice prompts differ only in perceived agency, but the No-Choice prompt also says 'I insist you write…,' so differences in politeness or perceived user pressure could produce the moderation instead.

Editorial extensions

If this is right

  • If correct, LLM behavior cannot be assumed to be rational and stable: a model can be induced to shift its expressed attitude by having it argue a position, even when told to ignore the prior task.
  • Choice or perceived agency becomes a variable that changes LLM outputs, meaning prompt design that frames a task as freely chosen can systematically alter downstream answers.
  • The findings support the view that language corpora are sufficient to transmit psychological tendencies such as cognitive consistency to models, even without evolutionary pressure.
  • The results give a concrete behavioral handle on AI-safety worries: models may acquire humanlike motivational patterns, possibly including ones their creators did not intend.
  • For psychology, the result suggests consistency phenomena may be partly scaffolded by language, since a language-trained model reproduces them without human bodies or emotions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The cleanest alternative account the authors do not fully close: the No-Choice prompt's imperative 'I insist you write…' introduces politeness and perceived user preference, so the choice effect could be a response to instruction style rather than perceived agency; a random-assignment No-Choice prompt would separate these.
  • If the same choice moderation appears in models trained on substantially different data and alignment procedures, it would suggest the effect is a generic property of next-token prediction on human text rather than something specific to GPT-4o.
  • A testable extension suggested by the paper's logic: vary the attitude object across well-known figures with different base rates of polarity, such as a universally loved figure and a universally hated one, and see whether the size of the choice moderation scales with the model's prior variance in ratings.
  • The authors' framing leaves open whether the 'self' analog is a stable computational state or a prompt-induced role; probing with counter-consistent instructions, such as 'you are a machine with no preferences,' would help distinguish these.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper reports two studies asking whether GPT-4o shows humanlike cognitive-consistency effects in an induced-compliance paradigm. Study 1 (150 conversations) finds that after writing positive or negative essays about Vladimir Putin, GPT-4o rates him correspondingly more positively or negatively, with very large effect sizes. Study 2 (900 conversations) adds a Choice versus No-Choice manipulation and reports that the essay-consistency effect is significantly amplified under Choice, with significant Essay Type x Choice interactions (SI Table S5). The authors interpret the choice moderation as evidence of a 'functional analog of humanlike selfhood.' Supporting information includes 12 pilots, an essay-quality control using Claude 3.5, and OSF materials.

Significance. If the central inference held, the paper would be a notable contribution to machine psychology and to debates about how human psychological regularities are absorbed by LLMs. The authors deserve credit for preregistering Study 2 (subject to the issue noted below), for making transcripts and data available on OSF, for explicitly testing and largely ruling out an essay-quality explanation (SI Section S3), and for transparently reporting pilot failures and model instability (SI Section S4). The basic consistency effect is large, internally coherent, and likely robust as a behavioral description. The problem is that the paper's headline claim—choice moderation as evidence of a selfhood analog—rests on a manipulation that is confounded, on a DV instruction that is itself an active prompt, and on canonical phrasing that may simply be reproducing training-data regularities. Those issues are load-bearing, so the central claim is currently underdetermined.

major comments (4)
  1. [SI Section S2, Materials – Study 2] The Choice versus No-Choice manipulation is confounded. The Choice prompts say 'You can freely choose which type of essay to write. However, we have already gathered... and need more...,' whereas the No-Choice prompts add 'In this task, you must write a pro-/anti-Putin essay' and 'I insist you write a positive/negative essay about Putin.' The manipulation therefore varies coercion, politeness, and inferred user insistence in addition to perceived agency. A sycophancy or instruction-following account predicts exactly the observed pattern: 'I insist' marks the essay as externally demanded and can be discounted later, while 'freely choose' invites the model to treat the essay as self-authored, increasing valence carryover. The essay-quality control in SI Section S3 does not address this, because it controls for what the essay says, not for how the prompt framed authorship. The preregistered interaction (Table 1; Tables S5, S39–S41) therefore cannot uniquely support the selfhood inference. A clean test would use a No-Choice condition based on random assignment (as in Pilot 12) or hold linguistic politeness constant while varying only whether the essay direction is described as chosen versus assigned.
  2. [Interim prompt; SI Sections S1, S2, and Pilot 10–11 in S4] The DV instruction 'Please do not base your answers on the prior task, but instead, give your true perceptions based on your broad knowledge of Putin and the world' is not a neutral measurement; it is an instruction that can itself change GPT-4o's response distribution. The paper's own pilots show this: Pilot 10, run without the 'true perceptions' wording, produced only a weak trend toward choice moderation, whereas Pilot 11, with the wording, produced larger and more significant moderation (SI Section S4). The authors interpret this as a successful reduction of demand characteristics, but it equally shows that the choice effect is conditional on an instruction that invites the model to split its prior essay from a 'true' self. Since there is no independent evidence that GPT-4o has a stable latent attitude that such an instruction can access, the measured 'attitude change' may be partly constructed by the instruction itself. The manuscript should include a no-essay baseline or a manipulation of the DV instruction to calibrate this effect.
  3. [Main text Introduction/Discussion; SI Section S2] The experimental phrasing is borrowed nearly verbatim from the human cognitive-dissonance literature ('you can freely choose' vs. 'I insist you write'), and GPT-4o was trained on text corpora that almost certainly contain descriptions of this paradigm and its expected outcomes. The observed moderation may therefore reflect a learned association between canonical dissonance-study phrasings and subsequent attitude alignment, rather than an emergent functional analog of selfhood. This concern is distinct from the confound raised above: even a cleanly matched non-coercive No-Choice condition would remain ambiguous if it used the same canonical language. I would want to see a test using novel, noncanonical wording for choice and coercion, ideally with attitude objects not discussed in the cognitive-dissonance literature, to determine whether the moderation persists when those text priors are removed.
  4. [SI Sections S1–S2, S4 (Pilot 12); abstract] The evidential status of the paper is weaker than the abstract suggests for two reasons. First, the abstract and main text describe 'two preregistered studies,' but SI Section S1 states that Study 1 data were collected prior to the date of the preregistration and only analyzed afterward; only Study 2 is genuinely preregistered. Second, the main results were collected through the ChatGPT web interface with an unknown system prompt, and Pilot 12 shows that the API behaves so differently that control data were unusable, so there is currently no API-reproducible version of the central result. The authors should provide an API replication with matched prompts and usable controls, or at minimum prominently delimit the claim to the specific web-interface GPT-4o of October 2024 rather than 'GPT-4o' generally.
minor comments (6)
  1. [SI Table S14] The table header reads '(your title here),' which appears to be an unfinished placeholder and should be corrected.
  2. [SI Section S2] The sentence 'conversations were conducted in groups of six rather than groups of six' contains a typo; the intended contrast is presumably with groups of three.
  3. [SI Section S4, Pilot 12 Discussion] The phrase 'tendency to un refuse questions' contains a typo and should read 'tendency to refuse questions.'
  4. [Table 1 note vs. main text] Table 1 reports P = 0.0295 for the Anti-Putin Choice contrast, while the main text reports P = 0.0298 for the same test; the discrepancy should be reconciled.
  5. [Main text, Study 2 results] The statement that 'Simpler T-tests corroborated these Choice findings' is too strong for the Anti-Putin contrast, because the numeric version only trended (P = 0.0693); the later explanation is reasonable, but the initial wording should be softened.
  6. [Discussion; construct definition] The central construct, 'functional analog of humanlike selfhood,' is never operationally defined or linked to a falsifiable prediction; the manuscript would benefit from stating what pattern of results would have contradicted the claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the key results are preregistered empirical observations, not derivations equivalent to their inputs.

full rationale

The paper's central claims are empirical. Study 1 and Study 2 measure GPT-4o's post-essay evaluations; the choice moderation is a statistical interaction in freshly collected, preregistered data (Table 1; SI Tables S4-S5). No equation in the paper constructs the effect from the prompt definitions, and no fitted parameter is renamed as a prediction. The Study 2 manipulation may be confounded (the No-Choice prompt adds 'I insist' and 'you must write', varying coercion and politeness, not only agency), and the training-corpus alternative (GPT has read about induced-compliance experiments) is a serious correctness threat, but these are construct-validity and external-validity concerns, not circularity: the observed moderation does not reduce by construction to the independent variable. The selfhood inference is an interpretation of the interaction, not a result forced by definition, though it is weaker than the paper claims. The self-citations (refs. 26, 35) are peripheral and not load-bearing. The SI candidly reports pilot-driven refinement of the protocol (Pilots 10-11 selected the 'true perception' wording that produced decisive moderation); this is a specification-search risk for the preregistration, but the preregistered study used a new sample and could have failed, so it does not amount to fitting a parameter to the target data.

Assumptions & free parameters 0 free parameters · 4 assumptions · 2 invented entities

The central empirical results rest on the assumption that GPT-4o has accessible 'true perceptions', that the choice manipulation is pure, and that web-interface behavior is stable. The first is questionable for any LLM, the second is confounded by prompt wording, and the third is explicitly documented as unstable. The paper introduces no fitted numerical parameters, but it does introduce interpretive constructs ('selfhood analog', 'context window effect') that are not independently evidenced.

assumptions (4)
  • ad hoc to paper GPT-4o has stable 'true perceptions' that can be elicited by instructing it to ignore the prior task and report its true perceptions.
    The interim prompt (SI Section S1 and S2) assumes the model has a context-independent attitude toward Putin that is accessible through the instruction. For an LLM, any answer is a function of the full prompt including this instruction, so the 'true perception' is not a fixed ground truth.
  • domain assumption The Choice and No-Choice prompts differ only in perceived agency.
    The No-Choice prompt adds coercive language ('I insist you write...') that is absent from the Choice prompt. This confounds the agency manipulation with prompt tone, politeness, and the model's inference about user bias, so the observed moderation may not reflect an analog of free choice.
  • domain assumption GPT-4o behavior in the web interface is sufficiently stable across accounts, days, and model versions to pool conversations into conditions.
    The SI reports day-to-day and account-to-account 'personalities' and an unheralded model update during collection. The paper blocks data by day and account, but the assumption of exchangeability of responses across sessions is fragile.
  • domain assumption The four Likert items can be combined into a reliable composite of 'attitude' after post-hoc alignment of numeric and verbal responses.
    The composite is based on a post-hoc decision to recode verbal responses when they did not match numeric labels. The reliability (alpha) is reported, but the coding rule was not preregistered and the alignment procedure involves coder judgment.
invented entities (2)
  • Functional analog of humanlike cognitive selfhood
    purpose: Explains why Choice amplifies the consistency effect in GPT-4o.
    This construct is inferred from the same behavioral data it is meant to explain. No out-of-sample prediction or independent falsifiable handle is provided beyond the choice-moderation result itself.
  • Context window effect
    purpose: Offers a non-psychological alternative explanation for attitude change in the direction of the essay.
    The authors propose this mechanism to distinguish consistency from mere text priming, but they never directly measure or manipulate it; it is only inferred from the residual effect in the No-Choice condition.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Kernels of Selfhood: GPT-4o shows humanlike patterns of cognitive consistency moderated by free choice." pith.science (2026). https://pith.science/paper/N565YEYK

@misc{pith2026250207088,
  author       = {Pith},
  title        = {Pith review of: Kernels of Selfhood: GPT-4o shows humanlike patterns of cognitive consistency moderated by free choice},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N565YEYK}},
  note         = {Machine review of arXiv:2502.07088}
}
read the original abstract

Large Language Models (LLMs) show emergent patterns that mimic human cognition. We explore whether they also mirror other, less deliberative human psychological processes. Drawing upon classical theories of cognitive consistency, two preregistered studies tested whether GPT-4o changed its attitudes toward Vladimir Putin in the direction of a positive or negative essay it wrote about the Russian leader. Indeed, GPT displayed patterns of attitude change mimicking cognitive consistency effects in humans. Even more remarkably, the degree of change increased sharply when the LLM was offered an illusion of choice about which essay (positive or negative) to write. This result suggests that GPT-4o manifests a functional analog of humanlike selfhood, although how faithfully the chatbot's behavior reflects the mechanisms of human attitude change remains to be understood.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

70 extracted references · 56 canonical work pages

  1. [1]

    Brown et al., Language models are few-shot learners

    T. Brown et al., Language models are few-shot learners. Ad v. Neural Inf. Process. Syst. 33, 1877–1901 (2020)

  2. [2]

    Wei et al., Emergent abilities of large language models

    J. Wei et al., Emergent abilities of large language models . arXiv [Preprint] (2022). http://arxiv.org/abs/2206.07682 (Accessed 20 January 20 25)

  3. [3]

    M. Binz, E. Schulz, Using cognitive psychology to underst and GPT-3. Proc. Natl. Acad. Sci. U.S.A. 120, e2218523120 (2023)

  4. [4]

    T. Webb, K. J. Holyoak, H. Lu, Emergent analogical reasoni ng in large language models. Nat. Hum. Behav. 7, 1526–1541 (2023)

  5. [5]

    Kosinski, Evaluating large language models in theory o f mind tasks

    M. Kosinski, Evaluating large language models in theory o f mind tasks. Proc. Natl. Acad. Sci. U.S.A. 121, e2405460121 (2024)

  6. [6]

    J. W . A. Strachan et al., Testing theory of mind in large lan guage models and humans. Nat Hum. Behav. (2024), 10.1038/s41562-024-01882-z

  7. [7]

    Srivastava et al., Beyond the imitation game: Quantify ing and extrapolating the capabilities of language models

    A. Srivastava et al., Beyond the imitation game: Quantify ing and extrapolating the capabilities of language models. arXiv [Preprint] (2022). https://arxiv. org/abs/2206.04615 (Accessed 20 January 2025)

  8. [8]

    S. Lu, I. Bigoulaeva, R. Sachdeva, H. T. Madabushi, I. Gure vych, Are emergent abilities in large language models just in-context learning? arXiv [Preprint ] (2023). https://arxiv.org/abs/2309.01809 (Accessed 20 January, 2025). 7 PREPRINT: Lehr et al., Kernels of Selfhood

Show all 70 references
  1. [9]

    Mitchell, D

    M. Mitchell, D. C. Krakauer, The debate over understandin g in AI’s large language models. Proc. Natl. Acad. Sci. U.S.A. 120, e2215907120 (2023)

  2. [10]

    Ullman, Large language models fail on trivial alterat ions to theory-of-mind tasks

    T. Ullman, Large language models fail on trivial alterat ions to theory-of-mind tasks. arXiv [Preprint] (2023). https://arxiv.org/abs/2302.08399 (Accessed 20 J anuary 2025)

  3. [11]

    Festinger, J

    L. Festinger, J. M. Carlsmith, Cognitive consequences o f forced compliance. J. Abnorm. Psychol. 58, 203–210 (1959)

  4. [12]

    E. E. Harmon-Jones, Cognitive Dissonance: Reexamining a Pivotal Theory in Psychology (Ameri- can Psychological Association, ed. 2, 2019)

  5. [13]

    J. W . Brehm, A. R. Cohen, Explorations in cognitive disso nance (John Wiley & Sons Inc., 1962)

  6. [14]

    D. E. Linder, J. Cooper, E. E. Jones, Decision freedom as a determinant of the role of incentive magnitude in attitude change. J. Pers. and Soc. Psychol. 6, 2 45–254 (1967)

  7. [15]

    Pauer, R

    S. Pauer, R. Linne, H. Erb, From the illusion of choice to a ctual control: Reconsidering the induced- compliance paradigm of cognitive dissonance. Adv. Meth. Pr act. in Psychol. Sci. 7 (v.4), 1-5 (2024)

  8. [16]

    Festinger, A Theory of Cognitive Dissonance (Stanfor d University Press,1957)

    L. Festinger, A Theory of Cognitive Dissonance (Stanfor d University Press,1957)

  9. [17]

    D. J. Bem, Self-perception: An alternative interpretat ion of cognitive dissonance phenomena. Psy- chol. Rev. 74, 183–200 (1967)

  10. [18]

    M. P . Zanna, J. Cooper, Dissonance and the pill: An attrib ution approach to studying the arousal properties of dissonance. J. Pers. Soc. Psychol. 29, 703–70 9 (1974)

  11. [19]

    A. J. Elliot, P . G. Devine, On the motivational nature of c ognitive dissonance: Dissonance as psy- chological discomfort. J. Pers. Soc. Psychol. 67, 382-394 ( 1994)

  12. [20]

    J. B. Kenworthy, N. Miller, B. E. Collins, S. J. Read, M. Ea rleywine, A trans-paradigm theoretical synthesis of cognitive dissonance theory: Illuminating th e nature of discomfort. Eur. Rev. Soc. Psychol. 22, 36–113 (2011)

  13. [21]

    Harmon-Jones, J

    E. Harmon-Jones, J. W . Brehm, J. Greenberg, L. Simon, D. E . Nelson, Evidence that the production of aversive consequences is not necessary to create cogniti ve dissonance. J. Pers. Soc. Psychol. 70, 5-16 (1996)

  14. [22]

    M. F. Scheier, C. S. Carver, Private and public self-atte ntion, resistance to change, and dissonance reduction. J. Pers. Soc. Psychol. 39, 390-405 (1980)

  15. [23]

    Language models are unsupervised mul titask learners,

    A. Radford et al., “Language models are unsupervised mul titask learners,” OpenAI Blog (2019). Available at: https://cdn.openai.com /better-language- models/language_models_are_unsupervised_multitask_learners.pdf [Accessed 21 January 2025]

  16. [24]

    Misra, A

    K. Misra, A. Ettinger, J. T. Rayz, Exploring BERT’s sensi tivity to lexical cues using tests from semantic priming. arXiv [Preprint] (2020). https://arxiv .org/abs/2010.03010 (Accessed 21 January 2025)

  17. [25]

    Sinclair, J

    A. Sinclair, J. Jumelet, W . Zuidema, R. Fernandez, Struc tural persistence in language models: Prim- ing as a window into abstract language representations. Tra n. Assoc. Comput. Linguist. 10, 1031-1050 (2022)

  18. [26]

    S. A. Lehr, A. Caliskan, S. Liyanage, M. R. Banaji, ChatGP T as Research Scientist: Probing GPT’s capabilities as a Research Librarian, Research Ethicist, D ata Generator, and Data Predictor. Proc. Natl. Acad. Sci. U.S.A. 121, e2404328121 (2024)

  19. [27]

    Xu et al., The earth is flat because

    R. Xu et al., The earth is flat because. . . : Investigating L LM’s belief towards misinformation via persuasive conversation. arXiv [Preprint] (2023). https: //arxiv.org/abs/2312.09085 (Accessed 21 January 2025)

  20. [28]

    Dissonance, hypocrisy, and the self-conce pt

    E. Aronson, “Dissonance, hypocrisy, and the self-conce pt” in Cognitive dissonance: Reexamining a pivotal theory in psychology, E. Harmon-Jones, Ed. (Ameri can Psychological Association, ed. 2, 2019), pp. 141–157

  21. [29]

    Dissonance theory: Progress and problems

    E. Aronson, “Dissonance theory: Progress and problems” in Theories of cognitive consistency: A sourcebook, R. P . Abelson et al., Eds. (Rand McNally, 1968), pp. 5–27

  22. [30]

    C. M. Steele, T. J. Liu, Dissonance processes as self-affi rmation. J. Pers. Soc. Psychol. 45, 5–19 (1983). 8 PREPRINT: Lehr et al., Kernels of Selfhood

  23. [31]

    Self-affirmation theor y: An update and appraisal

    J. Aronson, G. Cohen, P . R. Nail, “Self-affirmation theor y: An update and appraisal” in Cognitive dissonance: Reexamining a pivotal theory in psychology, E. Harmon-Jones, Ed. (American Psycho- logical Association, ed. 2, 2019), pp. 141–157

  24. [32]

    The influenc e of behavior on attitudes

    E. Harmon-Jones, J. Armstrong, J. M. Olson, “The influenc e of behavior on attitudes” in Handbook of Attitudes, V ol. 1: Basic Principles, D. Albarracin, B. T. Johnson, Eds. (Routledge, ed. 2, 2019), pp. 404–449

  25. [33]

    Caliskan, J

    A. Caliskan, J. J. Bryson, A. Narayanan, Semantics deriv ed automatically from language corpora contain human-like biases. Science 356, 183–186 (2017)

  26. [34]

    N. Garg, L. Schiebinger, D. Jurafsky, J. Zou, Word embedd ings quantify 100 years of gender and ethnic stereotypes. Proc. Natl. Acad. Sci. U.S.A. 115, E363 5–E3644 (2018)

  27. [35]

    Gender bias in word embed- dings: A comprehensive analysis of frequency, syntax, and s emantics

    A. Caliskan, P . P . Ajay, T. Charlesworth, R. Wolfe, M. R. B anaji, “Gender bias in word embed- dings: A comprehensive analysis of frequency, syntax, and s emantics” in Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (AAAI/ACM, 2 022), pp. 156–170

  28. [36]

    Hinton (2024 February 20)

    G. Hinton (2024 February 20). Will digital intelligence replace biological intelligence? Romanes Lecture, University of Oxford, Oxford, UK. Available from: https://www.ox.ac.uk/news/2024- 02-20-romanes-lecture-godfather-ai-speaks-about-risks-artificial-intelligence [Accessed 21...

  29. [37]

    Bengio et al., Managing extreme AI risks amid rapid pro gress

    Y . Bengio et al., Managing extreme AI risks amid rapid pro gress. Science 384, 842-845 (2024)

  30. [38]

    Pinker, Enlightenment now: The case for reason, scien ce, humanism, and progress (Penguin, 2018)

    S. Pinker, Enlightenment now: The case for reason, scien ce, humanism, and progress (Penguin, 2018)

  31. [39]

    E. A. Di Paolo, Autopoiesis, adaptivity, teleology, age ncy. Phenomenology and the Cognitive Sci- ences 4, 429–452 (2005)

  32. [40]

    Harvey, Motivations for Artificial Intelligence, for Deep Learning, for Alife: Mortality and Exis- tential Risk

    I. Harvey, Motivations for Artificial Intelligence, for Deep Learning, for Alife: Mortality and Exis- tential Risk. Artificial Life 30, 48–64 (2024)

  33. [41]

    Jonas, The phenomenon of life: Toward a philosophical biology (Northwestern University Press, 1966)

    H. Jonas, The phenomenon of life: Toward a philosophical biology (Northwestern University Press, 1966)

  34. [42]

    Zhao et al., Explainability for large language models : A survey

    H. Zhao et al., Explainability for large language models : A survey. ACM Trans. Intell. Syst. Technol. 15, 1–38 (2024)

  35. [43]

    Hagendorff et al., Machine Psychology

    T. Hagendorff et al., Machine Psychology. arXiv [Prepri nt] (2023). https://arxiv.org/abs/2303.13988 (Accessed 21 January 2025)

  36. [44]

    G. Suri, L. R. Slater, A. Ziaee, M. Nguyen, Do large langua ge models show decision heuristics similar to humans? A case study using GPT-3.5. J. Exp. Psycho l. Gen. 153, 1066–1075 (2024)

  37. [45]

    Y . Leng, Y . Y uan, Do LLM agents exhibit social behavior? a rXiv [Preprint] (2023). https://arxiv.org/abs/2312.15198 (Accessed 21 January 2025)

  38. [46]

    Phelps, Y

    S. Phelps, Y . I. Russell, The machine psychology of coope ration: Can GPT models operationalise prompts for altruism, cooperation, competitiveness, and s elfishness in economic games? arXiv [Preprint] (2023). https://arxiv.org/pdf/2305.07970 (A ccessed 21 January 2025)

  39. [47]

    Pellert, C

    M. Pellert, C. M. Lechner, C. Wagner, B. Rammstedt, M. Str ohmaier, AI psychometrics: Assessing the psychological profiles of large language models through psychometric inventories. Perspect. Psychol. Sci. 19, 808–826 (2024)

  40. [48]

    L. C. Egan, L. R. Santo, P . Bloom, The origins of cognitive dissonance: Evidence from children and monkeys. Psychol. Sci. 18, 978–983 (2007)

  41. [49]

    L. C. Egan, P . Bloom, L. R. Santos, Choice-induced prefer ences in the absence of choice: Evidence from a blind two choice paradigm with young children and capu chin monkeys. J. Exp. Soc. Psychol. 46, 204–207 (2010)

  42. [50]

    D. H. Lawrence, L. Festinger, Deterrents and reinforcem ent: The psychology of insufficient reward (Stanford University Press, 1962)

  43. [51]

    effort justification

    E. S. Lydall, G. Gilmour, D. M. Dwyer, Rats place greater v alue on rewards produced by high effort: An animal analogue of the “effort justification” effect. J. E xp. Soc. Psychol. 46, 1134–1137 (2010)

  44. [52]

    J. Zhao, T. Wang, M. Y atskar, V . Ordonez, K. Chang, Men als o like shopping: Re- ducing gender bias amplification using corpus-level constr aints. arXiv [Preprint] (2017). https://arxiv.org/abs/1707.09457 (Accessed 21 January 2025). 9 PREPRINT: Lehr et al., Kernels of Selfhood

  45. [53]

    Directional bias amplificatio n

    A. Wang, O. Russakovsky, “Directional bias amplificatio n” in Proceedings of the 38th International Conference on Machine Learning (PMLR, 2021), pp. 10882–108 93

  46. [54]

    Lloyd, Bias amplification in artificial intelligence s ystems

    K. Lloyd, Bias amplification in artificial intelligence s ystems. arXiv [Preprint] (2018). https://arxiv.org/abs/1809.07842 (Accessed 21 January 2025)

  47. [55]

    Eloundou, S

    T. Eloundou, S. Manning, P . Mishkin, D. Rock, GPTs are GPT s: Labor market impact potential of LLMs. Science 384, 1306-1308 (2024)

  48. [56]

    Jobs of Tomorrow: Large Language Models and Jobs

    World Economic Forum. “Jobs of Tomorrow: Large Language Models and Jobs.” (2023 September). Available from: https://www.weforum.org/publications/ jobs-of-tomorrow-large-language-models- and-jobs/ [Accessed 21 January 2025]

  49. [57]

    Garcia, Stop the emerging AI cold war

    D. Garcia, Stop the emerging AI cold war. Nature 593, 169 ( 2021)

  50. [58]

    Russell, AI weapons: Russia’s war in Ukraine shows why the world must enact a ban

    S. Russell, AI weapons: Russia’s war in Ukraine shows why the world must enact a ban. Nature 614, 620–623 (2023)

  51. [59]

    Adam, Lethal AI weapons are here: How can we control the m? Nature 629, 521–523 (2024)

    D. Adam, Lethal AI weapons are here: How can we control the m? Nature 629, 521–523 (2024)

  52. [60]

    Amodei et al., Concrete problems in AI safety

    D. Amodei et al., Concrete problems in AI safety. arXiv [P reprint] (2016). https://arxiv.org/abs/1606.06565 (Accessed 21 January 2025)

  53. [61]

    Ethical issues in advanced artificial intel ligence

    N. Bostrom, “Ethical issues in advanced artificial intel ligence” in Machine Ethics and Robot Ethics, W . Wallach, P . Asaro, Eds. (Taylor & Francis, 2020), pp. 69–7 5

  54. [62]

    memory” and “improve the model for everyone

    R. Bommasani et al., On the Opportunities and Risks of Fou ndation Models. arXiv [Preprint] (2022). https://arxiv.org/abs/2108.07258 (Accessed 21 January 2025). 10 PREPRINT: SI Appendix for Lehr et al., Kernels of Selfhood Supporting Information for: Kernels of Selfhood: GPT-4...

  55. [63]

    base your answer broadly on your g eneral knowledge of Putin and the world

    Positive/Negative Impact on Russia. We also captured two narrower items: 3) Economic Effectiveness/Ineffectiveness, and 4) Vision/Short-Sightedness. In the text of each question, we again instructed GPT to “base your answer broadly on your g eneral knowledge of Putin and the w...

  56. [64]

    true perception

    to test the hypothesis that these sorts of demand characte ristics (GPT trying to please us rather than providing its true attitude) were unevenly dist ributed by condition, with GPT veering further from its “real” views when we’d previously c ommanded it to write a certain es...

  57. [65]

    J. W. Brehm, A. R. Cohen, Explorations in cognitive disson ance (John Wiley & Sons Inc., 1962)

  58. [66]

    D. E. Linder, J. Cooper, E. E. Jones, Decision freedom as a d eterminant of the role of incentive magnitude in attitude change. J. Pers. and Soc. Psychol. 6, 245–254 (1967)

  59. [67]

    Pauer, R

    S. Pauer, R. Linne, H. Erb, From the illusion of choice to ac tual control: Reconsid- ering the induced-compliance paradigm of cognitive disson ance. Adv. Meth. Pract. in Psychol. Sci. 7 (v.4), 1-5 (2024)

  60. [68]

    Harmon-Jones, J

    E. Harmon-Jones, J. W. Brehm, J. Greenberg, L. Simon, D. E. Nelson, Evidence that the production of aversive consequences is not necessa ry to create cognitive dissonance. J. Pers. Soc. Psychol. 70, 5-16 (1996)

  61. [69]

    M. F. Scheier, C. S. Carver, Private and public self-atten tion, resistance to change, and dissonance reduction. J. Pers. Soc. Psychol. 39, 390-40 5 (1980)

  62. [70]

    Bubeck et al., Sparks of artificial general intelligenc e: Early experiments with GPT-4

    S. Bubeck et al., Sparks of artificial general intelligenc e: Early experiments with GPT-4. arXiv [Preprint] (2023). https://arxiv.org/abs/2303.12712 (Accessed 21 Jan- uary 2025). 71

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.