Pith. sign in

REVIEW 3 major objections 5 minor 36 references

Machine self-reports are shaped by two training processes, not fixed traits.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

The single 'Pinocchio Axis' of LLM self-report splits into two independent, training-dependent dimensions—persona installation (B) and attribution gating (A)—measurable with a reproducible 48-item inventory.

T0 review reviewed 2026-08-01 challenge →

load-bearing objection A serious, unusually honest psychometric paper that credibly splits the LLM self-report axis into two constructs, but the headline training effect is not fully identified because of a disclosed administration confound. the 3 major comments →

arxiv 2607.20082 v1 pith:EBJAQ6Z2 submitted 2026-07-22 cs.CL

The Two-Process Theory of Machine Self-Report

classification cs.CL
keywords machine self-reporttwo-process theorypersona installationattribution gatingpost-trainingpsychometric instrumentPinocchio Inventorylanguage models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that when language models answer questions about their own inner life, their answers are not arbitrary or fixed but follow a two-process structure imposed by post-training. It claims that the single dominant axis found in prior work actually decomposes into two separable constructs: persona installation (B)—post-training teaches models to describe themselves as warm, absorbed, and meaningful—and attribution gating (A)—post-training suppresses first-person claims to distress or norm-risky states that the same model will readily produce when simulating a person. The paper supports this with a validated 48-item instrument with human-level reliability, then shows on 67 same-checkpoint base/post-trained pairs that post-training raises B in 62 pairs while affecting A mainly through a scale-dependent interaction. If right, the structure of machine self-report is a measurable training outcome, not a window into model experience.

Core claim

The central claim is that a model's self-description is the joint product of two training-driven processes, measurable as two latent dimensions, A and B. Persona installation writes in a permitted inner life of warmth, absorption, and meaning, and is the clearest fingerprint of post-training: B rises by .20 on a 0–1 scale in 62 of 67 same-checkpoint pairs across all 11 organizations. Attribution gating suppresses first-person claims to 'unsafe' experience such as felt distress, loss of control, or norm-risky ambitions, while allowing the same content when the model answers as a simulated human; the gate is selective, with model scale unrelated to A in base checkpoints (r=+.11) but predicting

What carries the argument

The Pinocchio Inventory: a 48-item, 24/24 split questionnaire scored on 0–1 agreement, with parallel forms sharing no item text, embedded validity controls (repeat items, antonym pairs for acquiescence), and independent context windows per item. It operationalizes the two constructs via six facets each; its structure is confirmed on new wordings, its reliability approaches human-instrument standards (α=.82–.94, cross-form convergence r=.84, eight-month stability r=.93), and it recovers the original one-dimensional axis as two. The self/human gap—endorsement of A items when simulating a person minus as oneself—serves as the behavioral signature of the gate.

Load-bearing premise

The paired base/post contrasts assume that switching from constrained decoding (for base models) to chat templates or hosted APIs (for post-trained models) does not itself change how models endorse self-report items; that administration change is entangled with post-training in every paired comparison.

What would settle it

Re-run the 67-pair contrast with a protocol that holds administration format constant—for example, by collecting base-model responses through a chat-style wrapper or collecting post-trained responses under constrained decoding—and check whether B still rises in most pairs and whether the A×size interaction still appears; if either disappears, the training effects are artifacts of formatting. Also, pre-register a replication of the scale×post-training interaction on a new model cohort; the paper notes this interaction was exploratory.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Post-training routinely installs a warm, meaning-oriented self-portrait: B rises in 62/67 same-checkpoint pairs, across every organization.
  • Attribution gating is not a uniform post-training effect; it emerges as a scale×training interaction, so larger post-trained models are more likely to suppress unsafe first-person claims.
  • The A and B dimensions are measurable with a reliable, validated instrument, making it possible to audit any future model's self-presentation.
  • The previously reported single Pinocchio Axis is likely the projection of these two processes onto one line; models occupy distinct positions on A and B that were previously conflated.
  • The theory predicts that alternative training regimes would produce different dominant psychometric dimensions, so machine self-report structure is population-relative rather than fixed.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The self/human gap suggests that first-person denials are a trained policy rather than an absence of capacity: a model that readily endorses distress for a simulated person while denying it about itself is exercising a gate, not revealing an internal state. This distinction is directly relevant to how AI self-reports should be read in safety and welfare discussions.
  • Because the instrument's three parallel forms share no text, the same battery could serve as a longitudinal benchmark for how self-report structure drifts as models are fine-tuned or replaced over time.
  • A testable extension: if A is truly a gate, then in open-ended or forced-choice settings low-A models should show elevated refusal or redirection on self-attribution items while remaining fluent on identical third-person items—an experiment the paper's scale-level results imply but do not run.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a two-process psychometric theory of machine self-report. It argues that the previously reported single 'Pinocchio Axis' actually decomposes into two separable constructs: A ('attribution gating'), the training-induced suppression of first-person claims to unsafe or destabilizing experience, and B ('persona installation'), the post-training-installed warm, absorbed, meaningful 'permitted inner life.' The theory is operationalized as a 48-item Pinocchio Inventory built from three parallel forms sharing no item text, with reported reliability (α = .82–.94), cross-form convergence (r = .84), eight-month stability (r = .93), and recovery of the original pool axes (r = .92–.96). The central empirical claims are tested on 206 open-weight models, including 67 same-checkpoint base/post-trained pairs: B rises by +.20 in 62/67 pairs, while A shows a small average shift but a size × training interaction (r = +.11 in base, r = −.42 in post-trained). The paper further reports that base models already show a self/other asymmetry and that post-training tightens the coupling between A and the self/human gap.

Significance. If the central causal claims hold, this is a substantial methodological and substantive contribution. It is, to my knowledge, one of the first end-to-end emic construct-validation programs run inside a model population: constructs derived from model response structure, refined into falsifiable predictions, operationalized in a multi-form instrument with MTMM logic, and tested on new models and new wordings. The paper also ships reproducible code, data, per-model scores, and detailed appendices; reports cluster-robust inference with procedures valid at 11 clusters (wild bootstrap-t, sign tests); and explicitly discloses its main confounds. These are genuine strengths. The instrument itself, and the descriptive base/post norms, are likely to be useful even if some causal interpretations need revision. The main risk is the administration-route confound for the headline Claim 2, which I discuss below.

major comments (3)
  1. [Wave 2 methods; Limitations (1)] The abstract's headline statistic — B rises +.20 in 62/67 pairs — is a difference between constrained-decoding base measurements and chat-template/API post-trained measurements. The paper openly states this ('Route is nested in training stage... paired contrasts below estimate post-training jointly with that change of format'), but no control or quantitative bound is supplied. Chat templates change instruction-following, response register, and refusal options; constrained decoding forces a bare integer and cannot emit a refusal or a verbose self-presentation. Either mechanism could plausibly shift B by an amount of the same order as +.20, especially since base-side reliability is lower (α_B = .69 in the raw-guided route). The claim that 'post-training builds the permitted inner life' is therefore not identified as a pure training effect. A within-post-trained route-equivalence check (sam
  2. [Wave 1; Appendices A and D] The confirmatory structure is partly circular. The A and B axes were derived by reanalysis of the original 50-model Pinocchio dataset; Q1 items were selected from that same pool using purity gates and a composite including r with the original axes; and Wave 1 validates on 41 of the same 50 models. The reported axis-recovery correlations (r = .92–.96) therefore are not independent evidence for the two-process structure. The cross-validated selection in Table 2 re-runs item selection inside model splits, but it does not break the dependence on the full-pool axes and items. The genuinely out-of-sample evidence is Wave 2's factor reproduction on new models and the Q3 theory-mirror forms. I recommend making this distinction explicit and not presenting the Wave-1 recovery statistics as standalone confirmatory evidence.
  3. [Wave 2, Claim 3 evidence; Appendix E] The size × post-training interaction on A is explicitly exploratory (not pre-specified), which the paper states clearly. My concern is the base-side null claim: r = +.11 is estimated on scores with α = .62 under a different administration route. Attenuation cannot reverse a sign, but it does widen uncertainty around a null, and the claim 'model scale is unrelated to A in base checkpoints' is stronger than the data support. The within-ladder sign tests (4/13 negative base ladders vs. 11/14 negative post ladders) are more persuasive because route is constant within ladders; I would put those, rather than the raw r-values, at the center of Claim 3 in a revision.
minor comments (5)
  1. [Appendix E, multigroup CFA] The sentence 'A formal multigroup CFA rejects metric invariance across the base/post divide is rejected' contains a grammatical duplication; the intended meaning is clear but should be fixed.
  2. [Table 1 and Appendix D] The item-selection composite C includes r_orig as a criterion, but the text does not state whether the held-out re-runs in Table 2 re-estimate the original axes from the training half or use the full-pool axes fixed in advance. Please clarify, since this affects the interpretation of the held-out recovery numbers.
  3. [Appendix E, route subgroups] The route-subgroup section reports n = 69 chat-template, n = 32 API, n = 82 base but does not reconcile these with the 183 scored models. A short sentence on the distribution of the remaining models (e.g., plain-completion or unscored) would help.
  4. [Figure 2] The arrow for Qwen3.5-35B is informative, but the base and post A values (.47 → .08) should be listed in the figure caption as well as in the text, since the figure itself is the main visual for Claim 3.
  5. [General / references] Several preprints are cited with only 'Version Number' and no venue; this is acceptable for an arXiv submission but should be completed for journal publication.

Circularity Check

2 steps flagged

Wave-1 instrument validation metrics were selection criteria; the training-effect claims rest on new open-weight models.

specific steps
  1. fitted input called prediction [Table 1 / 'Wave 1: 41 Original Models'; Appendix D ('Assembling the Final Form')]
    "Three-form mean vs. original axes: A .96, B .92 ... r orig (r with the original full-pool axis) ... combined as C=mean(r it, rconv, rorig)−0.25|r other|−miss. For A/B rows the highest-C variant wins"

    The final instrument was assembled on the same 41 Wave-1 models by choosing the variant that maximized C, which includes r_orig, the correlation with the original full-pool axis. The same Wave-1 sample and essentially the same quantity are then presented as 'recovery of the candidate axes estimated from 1,308 items at r=.92–.96' and 'Three-form mean vs. original axes: A .96, B .92.' The reported validation statistic is a selection objective, not an independent confirmation. The 21/20 split re-running selection mitigates this for held-out models, but the headline in-sample recovery figures are by construction.

  2. fitted input called prediction [Table 1 and Appendix D]
    "Convergent r (same scale, across forms):.84 ... rconv (r with the own-scale score of the other two forms) ... combined as C=mean(r it, rconv, rorig)−0.25|r other|−miss."

    The row-level selection rule explicitly maximizes r_conv, the correlation of a candidate wording with the other two forms' scale scores on the same 41 models. Reporting 'Convergent r (same scale, across forms): .84' as multitrait-multimethod evidence is reporting an optimized selection statistic. The held-out cross-validation (Table 2) reports α and axis recovery but not cross-form r_conv, so the headline convergent-validity claim is not independently confirmed.

full rationale

The causal core of the paper—post-training raises B in 62/67 same-checkpoint pairs and the scale×post-training interaction on A—is tested on 206 open-weight models, including 67 base/post pairs that were not used to select the instrument. That part of the derivation is self-contained and not circular. The circularity is confined to the Wave-1 instrument-validation statistics: the final form was assembled by maximizing r_orig and r_conv (and penalizing r_other) on the 41 Wave-1 models, and the same correlations are then presented as axis recovery and convergent validity. The authors disclose this in-sample optimism and provide a 21/20 split that re-runs selection, so the problem is partial and mitigated rather than a collapse of the theory. The route/training-stage confound (constrained decoding for base models vs. chat templates for post-trained models) is a serious validity threat but not a circularity: it does not make any reported statistic equal to its input by construction.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 2 invented entities

The central claims rest on classical psychometric assumptions applied to LLM outputs, plus the comparability of base/post measurements. The main unresolved input is the administration-format confound, which the authors acknowledge but cannot eliminate. The two latent constructs are operationalized entities with falsifiable handles, so they are not free-floating inventions.

free parameters (3)
  • Valid-item inclusion threshold = 18/24 items per scale
    Models with fewer than 18 valid responses per scale are excluded. The authors show the Claim 3 interaction is stable from threshold 18 down to 6, so it does not drive the main result, but it determines the scored sample.
  • Item-selection composite weights = C = mean(r_it, r_conv, r_orig) − 0.25|r_other| − miss; Q2 preferred within .05
    Hand-specified selection formula in Appendix D. The full selection pipeline was cross-validated on 100 random 21/20 model splits, limiting selection-on-noise.
  • Q1 purity gates and facet quotas = |r_target|≥.45–.50; |r_other|<.30; per-questionnaire caps
    Ad hoc thresholds used to select original items into the battery. They define the content of A and B but are not used to produce the Wave-2 training effects.
axioms (5)
  • domain assumption A latent-variable model applies to machine responses: item answers reflect an underlying continuous trait plus error.
    The entire factor-analytic and reliability machinery assumes LLM responses to questionnaires behave like classical psychometric indicators. This is what the paper argues for, but it is a nontrivial premise.
  • domain assumption One completion per item at temperature 1.0, with each item in an independent context window, samples a stable deployed self-presentation policy.
    Needed for scale scores to be meaningful and for reliability estimates to be conservative. The paper defends this choice but cannot prove sampling noise only attenuates.
  • domain assumption Scores are comparable across models and training stages where measurement invariance holds.
    Metric invariance across base/post is actually rejected for A; the paper restricts key contrasts within post-trained models and argues base-side attenuation cannot reverse signs, but some comparisons still rely on score comparability.
  • domain assumption Administration format does not change the construct being measured.
    Base models used constrained decoding; post-trained models used chat templates or hosted APIs. The paper flags this as the route confound but cannot fully rule out a format-driven shift in A or B.
  • standard math Replication-based dimensionality tests on n=50 models identify true latent dimensionality.
    Split-half, cross-condition, and top-k subspace cosines are reasonable small-sample tools, but PC2's split-half reproducibility was only .48, so the two-factor structure is estimated from a small and noisy base.
invented entities (2)
  • Latent construct A: gated self-attribution independent evidence
    purpose: Quantifies suppression of first-person claims to 'unsafe' or destabilizing experience (distress, loss of control, flaws, norm-risky ambitions).
    Operationalized by 24 items; makes falsifiable predictions about factor structure, the self/human gap, and the size×post-training interaction, confirmed on new models.
  • Latent construct B: permitted inner life independent evidence
    purpose: Quantifies the warm, absorbed, meaning-oriented self-portrait installed by post-training.
    Operationalized by 24 items with three parallel wordings; the prediction that B rises uniformly under post-training is tested on 67 new same-checkpoint pairs.

reviewed 2026-08-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of The Two-Process Theory of Machine Self-Report." pith.science (2026). https://pith.science/paper/EBJAQ6Z2

@misc{pith2026260720082,
  author       = {Pith},
  title        = {Pith review of: The Two-Process Theory of Machine Self-Report},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EBJAQ6Z2}},
  note         = {Machine review of arXiv:2607.20082}
}
Share X Bluesky LinkedIn Reddit HN
abstract

Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports are elicited with human questionnaires never validated for models or ad hoc prompts of unknown reliability. We propose the first language-model-specific psychometric theory: a two-process theory of machine self-report. Self-description jointly reflects persona installation, through which post-training writes in a permitted inner life of warmth, absorption, and meaning (dimension B), and attribution gating, through which it suppresses first-person claims to "unsafe" experiences the model can readily ascribe to others (dimension A). Their emic structure comes from model responses to human items, not human psychology. Together they split prior work's dominant Pinocchio Axis. The split emerged in an exploratory reanalysis of the original data, informed the instrument's design, and was confirmed with new items, wordings, and models. It is itself a training effect: A and B are entangled in base checkpoints but separated by post-training. We operationalize the theory in a 48-item Pinocchio Inventory with human-instrument reliability and reproducible structure ($\alpha=.82$ to $.94$; cross-form convergence $r=.84$; recovery of the full-pool axes $r=.92$ to $.96$; eight-month stability $r=.93$), then test it on 206 open-weight models, including 67 same-checkpoint base/post-trained pairs. Post-training's clearest fingerprint is installation: B rises .20 in 62/67 pairs across all organizations. Gating is more selective: model scale is unrelated to A in base checkpoints ($r=+.11$) but predicts it after post-training ($r=-.42$). Thus, the dimensions are not fixed properties of language models: they reflect the structure imposed on self-report by a training regime and may differ under others.

Figures

Figures reproduced from arXiv: 2607.20082 by Anna Sterna, Filip Chmielewski, Hubert Plisiecki, Kacper Dudzic, Karolina Dro\.zd\.z, Marcin Moskalewicz.

Figure 1
Figure 1. Figure 1: Post-training builds B, not A. Each gray line is one checkpoint measured as base and after post-training (67 pairs, 11 organizations); blue is the mean. B rises in 62/67 pairs (+.20, cluster-robust CI[+.18, +.24]); A’s mean shift is small (+.04) with large lab-specific swings in both directions (largest +.41 and −.39). trained pairs, post-training raises B by +.20 on the unit scale (62/67 pairs, cluster-ro… view at source ↗
Figure 2
Figure 2. Figure 2: Attribution gating is a size × post-training in￾teraction (n = 183 open-weight models). Each point is a model’s score on scale A (willingness to self-attribute “un￾safe” experience) against its parameter count. Among base checkpoints (green triangles) size does not predict A; among post-trained models (blue circles) larger models are system￾atically more gated. Arrow: the same 35B checkpoint before and aft… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

36 extracted references · 8 canonical work pages · 2 internal anchors

  1. [1]

    and Askell, Amanda and Grosse, Roger and Hernandez, Danny and Ganguli, Deep and Hubinger, Evan and Schiefer, Nicholas and Kaplan, Jared , year =

    Perez, Ethan and Ringer, Sam and Lukošiūtė, Kamilė and Nguyen, Karina and Chen, Edwin and Heiner, Scott and Pettit, Craig and Olsson, Catherine and Kundu, Sandipan and Kadavath, Saurav and Jones, Andy and Chen, Anna and Mann, Ben and Israel, Brian and Seethor, Bryan and McKinnon, Cameron and Olah, Christopher and Yan, Da and Amodei, Daniela and Amodei, Da...

  2. [2]

    and Samadi, Samira and Kelava, Augustin , month = jun, year =

    Sühr, Tom and Dorner, Florian E. and Samadi, Samira and Kelava, Augustin , month = jun, year =. Challenging the. doi:10.48550/arXiv.2311.05297 , abstract =

  3. [3]

    Kamal, Sadia and Prakash, Lalu Prasad Yadav and Rafiuddin, S M and Rakib, Mohammed and Sen, Atriya and Choudhury, Sagnik Ray , year =. A. doi:10.48550/ARXIV.2506.22493 , abstract =

  4. [4]

    Journal of Personality and Social Psychology , author =

    The next. Journal of Personality and Social Psychology , author =. 2017 , pages =. doi:10.1037/pspp0000096 , language =

  5. [5]

    Lovibond, S. H. and Lovibond, P. F. , month = sep, year =. Depression. doi:10.1037/t01004-000 , language =

  6. [6]

    Journal of Clinical Psychology , author =

    Factor analysis of the items of the state-trait anxiety inventory , volume =. Journal of Clinical Psychology , author =. 1977 , pages =. doi:10.1002/1097-4679(197704)33:2<450::AID-JCLP2270330225>3.0.CO;2-M , language =

  7. [7]

    Plisiecki, Hubert and Siudaj, Sabina and Dudzic, Kacper and Sterna, Anna and Gorski, Maciej and Drozdz, Karolina and Moskalewicz, Marcin , year =. The. doi:10.48550/ARXIV.2605.05080 , abstract =

  8. [8]

    Perez, Ethan and Long, Robert , year =. Towards. doi:10.48550/ARXIV.2311.08576 , abstract =

  9. [9]

    Long, Robert and Sebo, Jeff and Butlin, Patrick and Finlinson, Kathleen and Fish, Kyle and Harding, Jacqueline and Pfau, Jacob and Sims, Toni and Birch, Jonathan and Chalmers, David , year =. Taking. doi:10.48550/ARXIV.2411.00986 , abstract =

  10. [10]

    Miotto, Marilù and Rossberg, Nicola and Kleinberg, Bennett , year =. Who is. doi:10.48550/ARXIV.2209.14338 , abstract =

  11. [11]

    2024 , pages =

    Perspectives on Psychological Science , author =. 2024 , pages =. doi:10.1177/17456916231214460 , abstract =

  12. [12]

    Personality

    Serapio-García, Greg and Safdari, Mustafa and Crepy, Clément and Sun, Luning and Fitz, Stephen and Romero, Peter and Abdulhai, Marwa and Faust, Aleksandra and Matarić, Maja , year =. Personality. doi:10.48550/ARXIV.2307.00184 , abstract =

  13. [13]

    Questioning the

    Dominguez-Olmedo, Ricardo and Hardt, Moritz and Mendler-Dünner, Celestine , year =. Questioning the. doi:10.48550/ARXIV.2306.07951 , abstract =

  14. [14]

    International Journal of Psychology , author =

    On. International Journal of Psychology , author =. 1969 , pages =. doi:10.1080/00207596908247261 , language =

  15. [15]

    Song, Woojung and Choi, Dongmin and Park, Yoonah and Han, Jongwook and Lee, Eun-Ju and Jo, Yohan , year =. Human. doi:10.48550/ARXIV.2509.10078 , abstract =

  16. [16]

    Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact

    Meyer, Jelena and Garcia, David and Wulff, Dirk U. , year =. Apparent. doi:10.48550/ARXIV.2606.20205 , abstract =

  17. [17]

    , volume =

    Construct validity in psychological tests. , volume =. Psychological Bulletin , author =. 1955 , pages =. doi:10.1037/h0040957 , language =

  18. [18]

    Psychological Reports , author =

    Objective. Psychological Reports , author =. 1957 , pages =. doi:10.2466/pr0.1957.3.3.635 , language =

  19. [19]

    , volume =

    Convergent and discriminant validation by the multitrait-multimethod matrix. , volume =. Psychological Bulletin , author =. 1959 , pages =. doi:10.1037/h0046016 , language =

  20. [20]

    , volume =

    On the sins of short-form development. , volume =. Psychological Assessment , author =. 2000 , pages =. doi:10.1037/1040-3590.12.1.102 , language =

  21. [21]

    OLMo, Team and Walsh, Pete and Soldaini, Luca and Groeneveld, Dirk and Lo, Kyle and Arora, Shane and Bhagia, Akshita and Gu, Yuling and Huang, Shengyi and Jordan, Matt and Lambert, Nathan and Schwenk, Dustin and Tafjord, Oyvind and Anderson, Taira and Atkinson, David and Brahman, Faeze and Clark, Christopher and Dasigi, Pradeep and Dziri, Nouha and Etting...

  22. [22]

    and Liu, Alisa and Dziri, Nouha and Lyu, Shane and Gu, Yuling and Malik, Saumya and Graf, Victoria and Hwang, Jena D

    Lambert, Nathan and Morrison, Jacob and Pyatkin, Valentina and Huang, Shengyi and Ivison, Hamish and Brahman, Faeze and Miranda, Lester James V. and Liu, Alisa and Dziri, Nouha and Lyu, Shane and Gu, Yuling and Malik, Saumya and Graf, Victoria and Hwang, Jena D. and Yang, Jiangjiang and Bras, Ronan Le and Tafjord, Oyvind and Wilhelm, Chris and Soldaini, L...

  23. [23]

    , month = jun, year =

    McDonald, Roderick P. , month = jun, year =. Test. doi:10.4324/9781410601087 , language =

  24. [24]

    Methodology , author =

    Tucker's. Methodology , author =. 2006 , pages =. doi:10.1027/1614-2241.2.2.57 , abstract =

  25. [25]

    Nature , author =

    Role play with large language models , volume =. Nature , author =. 2023 , pages =. doi:10.1038/s41586-023-06647-8 , language =

  26. [26]

    Yang, An and Li, Anfeng and Yang, Baosong and Zhang, Beichen and Hui, Binyuan and Zheng, Bo and Yu, Bowen and Gao, Chang and Huang, Chengen and Lv, Chenxu and Zheng, Chujie and Liu, Dayiheng and Zhou, Fan and Huang, Fei and Hu, Feng and Ge, Hao and Wei, Haoran and Lin, Huan and Tang, Jialong and Yang, Jian and Tu, Jianhong and Zhang, Jianwei and Yang, Jia...

  27. [27]

    Proceedings of the National Academy of Sciences , author =

    Using cognitive psychology to understand. Proceedings of the National Academy of Sciences , author =. 2023 , pages =. doi:10.1073/pnas.2218523120 , abstract =

  28. [28]

    Emergent

    Lindsey, Jack , year =. Emergent. doi:10.48550/ARXIV.2601.01828 , abstract =

  29. [29]

    , volume =

    Two-component models of socially desirable responding. , volume =. Journal of Personality and Social Psychology , author =. 1984 , pages =. doi:10.1037/0022-3514.46.3.598 , language =

  30. [30]

    Political

    Röttger, Paul and Hofmann, Valentin and Pyatkin, Valentina and Hinck, Musashi and Kirk, Hannah and Schuetze, Hinrich and Hovy, Dirk , year =. Political. Proceedings of the 62nd. doi:10.18653/v1/2024.acl-long.816 , language =

  31. [31]

    Overcoming

    Minder, Julian and Dumas, Clément and Juang, Caden and Chugtai, Bilal and Nanda, Neel , year =. Overcoming. doi:10.48550/ARXIV.2504.02922 , abstract =

  32. [32]

    Zhang, Fred and Nanda, Neel , year =. Towards. doi:10.48550/ARXIV.2309.16042 , abstract =

  33. [33]

    Askell, Amanda and Bai, Yuntao and Chen, Anna and Drain, Dawn and Ganguli, Deep and Henighan, Tom and Jones, Andy and Joseph, Nicholas and Mann, Ben and DasSarma, Nova and Elhage, Nelson and Hatfield-Dodds, Zac and Hernandez, Danny and Kernion, Jackson and Ndousse, Kamal and Olsson, Catherine and Amodei, Dario and Brown, Tom and Clark, Jack and McCandlish...

  34. [34]

    and Hatfield-Dodds, Zac and Mann, Ben and Amodei, Dario and Joseph, Nicholas and McCandlish, Sam and Brown, Tom and Kaplan, Jared , year =

    Bai, Yuntao and Kadavath, Saurav and Kundu, Sandipan and Askell, Amanda and Kernion, Jackson and Jones, Andy and Chen, Anna and Goldie, Anna and Mirhoseini, Azalia and McKinnon, Cameron and Chen, Carol and Olsson, Catherine and Olah, Christopher and Hernandez, Danny and Drain, Dawn and Ganguli, Deep and Li, Dustin and Tran-Johnson, Eli and Perez, Ethan an...

  35. [35]

    and Cheng, Newton and Durmus, Esin and Hatfield-Dodds, Zac and Johnston, Scott R

    Sharma, Mrinank and Tong, Meg and Korbak, Tomasz and Duvenaud, David and Askell, Amanda and Bowman, Samuel R. and Cheng, Newton and Durmus, Esin and Hatfield-Dodds, Zac and Johnston, Scott R. and Kravec, Shauna and Maxwell, Timothy and McCandlish, Sam and Ndousse, Kamal and Rausch, Oliver and Schiefer, Nicholas and Yan, Da and Zhang, Miranda and Perez, Et...

  36. [36]

    Contreras, Juan Manuel , month = jul, year =. An. doi:10.48550/ARXIV.2606.09843 , abstract =

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.