Pith. sign in

REVIEW 3 major objections 6 minor 77 references

Persona Cartography: Charting Language Model Personality Traits in Weight Space

T0 review · 3 major / 6 minor · reviewed 2026-07-10 · grok-4.5

Pith's one-line read LLM personas can be steered as composable trait directions in weight space, using OCEAN LoRAs that scale, mix, and shift safety behaviour.

desk verdict Solid empirical toolkit paper: OCEAN LoRAs scale and compose across six models with real safety side-effects; measurement partly circular but not hollow, and the work is worth engaging. read the letter →

arxiv 2607.07916 v1 pith:H26P3Z55 submitted 2026-07-08 cs.AI cs.LG

classification cs.AIcs.LG
keywords LLMpersonasOCEANtraitsLoRAadaptersweight-spaceeditingpersonacontrolsycophancyfrustrationpsychometrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Language models do not just answer questions; they behave with stable styles of deference, energy, caution, and tone that shape how they generalise and how safe they are. This paper treats those styles as positions in a trait space, starting from the five OCEAN personality dimensions, and trains low-rank adapters that amplify or suppress each trait. Across six models from three families, each adapter moves its target trait largely monotonically with scale, combines approximately additively with other adapters, and leaves capability benchmarks intact at moderate strengths. The same axes also move held-out safety behaviours: neuroticism tracks multi-turn frustration, agreeableness tracks sycophancy and compliance trade-offs. An unsupervised pipeline then recovers four model-native factors—tone, initiative, didacticism, and epistemic caution—from free rollouts, showing that the method is not locked to human psychology. The practical claim is that persona control becomes a matter of learning, scaling, and composing trait directions in weight space rather than one-off prompting or full retraining.

What carries the argument

Composable OCEAN trait LoRAs: low-rank adapters trained by constitution-guided DPO plus a lighter SFT stage on self-interaction transcripts, then scaled and linearly summed in weight space so that each adapter acts as a continuous, signed direction of persona change.

What would settle it

If, on held-out free-form multi-turn rollouts scored by independent human raters who never saw the training constitutions, scaling an openness amplifier failed to raise openness while leaving non-target traits and capability scores essentially unchanged, the central claim that the adapters are targeted trait directions would fail.

Watch

Extended reading notes

Core claim

Across six models (4B–32B, three families), constitution-trained OCEAN LoRA adapters move their target behavioural traits largely monotonically with scale, compose approximately additively into mixed personas, preserve MMLU and related capabilities at moderate scales, and shift held-out safety-relevant behaviours such as frustration and sycophancy.

Load-bearing premise

That human OCEAN labels, the same constitutions used in training, and TRAIT-style multiple-choice plus judges built from those definitions isolate real generalisable behavioural axes rather than teacher style, verbosity, or other correlated surface artefacts.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes treating LLM personas as positions in a behavioural trait space, operationalised via OCEAN (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism). It trains rank-64 LoRA amplifiers and suppressors for each trait using constitution-guided DPO plus SFT distillation (Open Character Training), evaluates them with TRAIT logprob MCQs, a human-calibrated LLM-judge panel, and capability benchmarks (MMLU, GSM8K, TruthfulQA), and shows that adapters move target traits largely monotonically with scale, compose approximately additively, and preserve capabilities at moderate scales across six models (4B–32B, three families). Held-out safety evaluations link neuroticism to multi-turn frustration and agreeableness to sycophancy/compliance; an unsupervised factor analysis on rollouts recovers four TIDE factors (Tone, Initiative, Didacticism, Epistemic Caution), with partial LoRA control of Initiative.

Significance. If the results hold, the work supplies a practical, weight-space alternative to brittle prompting and layer-sensitive activation steering for persona control: lightweight, scalable, and composable trait adapters that transfer across model families and teachers, retain rank-1 compressibility, and move safety-relevant behaviours without task-specific training. Strengths include multi-pronged measurement (TRAIT, human-calibrated judges with reported Spearman ρ/MAE and Krippendorff α, Wilson/bootstrap CIs), neutral control adapters, two distillation teachers, six baselines, interaction residuals (Eq. form in Appendix F / Fig. 4c), amplifier×suppressor heatmaps, and explicit safety trade-off experiments (WildJailbreak, CoCoNot, sycophancy). The unsupervised TIDE pipeline is a useful step beyond human psychometrics. The paper is a solid bridge between personality measurement, model editing, and alignment practice.

major comments (3)
  1. [§2.1 Methods; Appendix B, C.2.1] §2.1 / Appendix B and C.2.1: Constitutions, judge rubrics, and system-prompt baselines are all generated from the same OCEAN-definition object. TRAIT is also an OCEAN instrument. Monotonicity and near-additivity (Figs. 2–4) therefore partly re-measure definition-aligned style. Held-out safety tasks and human judge calibration partially break this loop, but the main text should quantify how much TRAIT/judge movement remains after residualising out surface correlates (length, sentiment, first-person affect) and should state more sharply which claims rest on external vs. definition-internal evidence.
  2. [§3 Downstream Applications] §3 and Figs. 39–40: Neutral control adapters nearly double sycophancy (0.61 vs 0.33 baseline), raise WildJailbreak harmful compliance, and modestly dampen frustration. Distillation artefacts are therefore not negligible relative to trait effects. Safety claims that attribute shifts to OCEAN axes (neuroticism↔frustration, agreeableness↔sycophancy) need systematic control-subtracted effect sizes and, where possible, matched-verbosity or matched-preference-strength baselines so trait signal is separated from pipeline shift.
  3. [§5.1; Abstract; §2.3] §5.1 Limitations: The full stack (judges, multi-turn safety, composition residuals) is reported primarily on Llama-3.1-8B-Instruct; other models receive TRAIT+MMLU only. The abstract’s “across six models” claim for composability and safety transfer is stronger than the evaluation breadth supports. Either extend at least one safety task and one composition residual analysis to a second family (e.g. Gemma-3-27B-IT, already used for frustration) or narrow the abstract/intro wording to match what was fully measured.
minor comments (6)
  1. [§2.3 Scaling and Combining LoRAs] Negative scaling does not reliably invert traits for all adapters (Appendix E); this should be flagged in the main §2.3 invertibility paragraph rather than only in the appendix, since signed-axis language appears in the abstract.
  2. [§4 Unsupervised Persona Exploration] §4 / Appendix M.5: Initiative LoRAs are validated on the same forced-choice questionnaire used to discover TIDE factors. Note this circularity more prominently when claiming “partial success” at modulating unsupervised traits.
  3. [Fig. 4c; Appendix F] Fig. 4c and Appendix F: Conscientiousness is the clear residual outlier; a one-sentence mechanistic hypothesis (or explicit “unknown”) in the main text would help readers interpret non-additivity.
  4. [Appendix A.1.3; §2.3] Appendix A.1.3: The factor-space merge (√w A/B) introduces cross terms distinct from the elementwise ΔW composition in §2.3; a short clarifying sentence in the main methods would prevent conflating the two “composition” notions.
  5. [Appendix E] Several figure panels (e.g. stacked MMLU bars at extreme scales) are dense; ensuring colourblind-safe palettes and consistent axis ranges across Appendix E sweeps would improve readability.
  6. [§5.2 Related Work] Related work: Sun et al. (2025) personality-vector merging and Vu et al. (2026) PsychAdapter are discussed; a short table contrasting rank, training objective, and composition method would make the novelty claim sharper.

Circularity Check

2 steps flagged · score 3.0 of 10

Mostly empirical, not definitionally circular; residual circularity is measurement alignment (same OCEAN object for train and free-form judges) and unsupervised validation on the discovery questionnaire—not a forced derivation of the main claims.

  1. self definitional [Appendix C.2.1 Judge Rubrics; §2.1 Measuring trait expression]
    "Every judge prompt is built mechanically from a single canonical OCEAN-definition object that also drives the training constitutions and the system-prompt induction baseline. This guarantees that the trait being trained, the trait being scored, and the trait being read aloud as a system prompt all refer to the same construct: the same facets, the same adjectives, the same canonical voice examples."

    Free-form trait scores used to claim that adapters 'move the target trait' are generated from the identical facet catalogue and voice examples used to train those adapters. On the judge axis alone, success is partly definition-aligned style matching rather than an independent behavioural measurement. TRAIT MCQs and held-out safety tasks break full circularity for the paper's strongest claims, but judge-based monotonicity/composition figures (e.g. Fig. 2, Fig. 4) inherit this shared construct.

  2. fitted input called prediction [§4 Unsupervised Persona Exploration; Limitations §5.1]
    "Amplifier and suppressor LoRAs were trained for the first factor ranked by explained-variance, Initiative, using the constitutional approach of Section 2, and the questionnaire re-administered under each adapter. ... Validation of trait-induction was performed using the questionnaire which found the factors, rather than independent judges as discussed above."

    TIDE factors are defined by loadings on a forced-choice questionnaire; Initiative LoRA 'success' is then reported as shifts on that same questionnaire (Fig. 8). The validation quantity is the discovery instrument, so factor-score movement is not an independent prediction of a new behavioural measure. The paper acknowledges this; it weakens only the unsupervised control claim, not the supervised OCEAN+safety results.

full rationale

This is an empirical model-editing paper, not a first-principles derivation. The central claims (monotonic TRAIT/judge movement, approximate LoRA additivity, capability retention, held-out safety shifts) are experimental outcomes, not quantities fitted then re-predicted. TRAIT is an external MCQ instrument; sycophancy/CoCoNot/WildJailbreak/frustration protocols are held out of training; control adapters and multi-model TRAIT+MMLU sweeps provide independent checks. Two limited circularities remain: (1) free-form LLM judges are built from the same canonical OCEAN-definition object as the training constitutions, so judge-score monotonicity partly re-reads definition-aligned style (mitigated by TRAIT and human calibration, but not eliminated); (2) unsupervised TIDE factors are validated by re-scoring the same forced-choice questionnaire used to recover them, which the paper itself flags in Limitations. Neither step makes the main weight-space control results true by construction. No self-citation uniqueness theorem, no fitted parameter renamed as prediction of an equivalent quantity, and no ansatz smuggled as a theorem. Score 3 reflects real but partial measurement circularity, not a collapsed derivation.

Assumptions & free parameters 5 free parameters · 5 assumptions · 2 invented entities

The central claim is empirical and rests on standard ML tools plus the domain assumption that human OCEAN (and later TIDE) axes are useful coordinates for LLM behaviour. Free parameters are training and scaling knobs chosen by convention or prior character-training recipes, not fitted to the safety outcomes. Invented entities are mainly the recovered TIDE factors and the specific trait LoRAs as operational objects; independent evidence for TIDE is partial (cross-model loading, internal consistency) but validation of Initiative is partly circular.

free parameters (5)
  • LoRA rank and alpha
    Rank 64, α=128 applied to all attention and MLP matrices; default taken from Open Character Training, not derived.
  • SFT adapter shrink factor 0.25
    SFT LoRA weights multiplied by 0.25 before factor-space merge with DPO LoRA, following Maiya et al.; hand-chosen recipe weight.
  • DPO β and NLL coefficient
    β=0.1, NLL coefficient 0.1, lr 5e-5, one epoch; standard hyperparameters that affect how strongly traits embed.
  • LoRA scale coefficients c
    Continuous scalar multipliers on ΔW used as the control knob; practical operating range |c|≲2 is empirical, not theoretically fixed.
  • Judge and TRAIT subsample sizes
    e.g. 240/299 seed prompts, 300 of 1000 TRAIT items, 100 MMLU items × 3 runs; design choices that affect score variance.
assumptions (5)
  • domain assumption Human Big-Five (OCEAN) trait structure is a useful initial basis for LLM behavioural dispositions.
    Stated as central insight and starting point in Abstract and §2; justified by human psychometrics literature, not proven for models.
  • domain assumption Constitution-conditioned teacher pairs plus DPO+SFT isolate the intended trait rather than only teacher style or verbosity.
    Core of §2.1 training; partially tested via neutral control adapters, which still shift safety metrics.
  • domain assumption Calibrated LLM judges and TRAIT logprob scores are valid measures of the same constructs the constitutions target.
    Appendix C.2 calibration against small human gold sets; TRAIT is human-designed (Limitation §5.1).
  • domain assumption Elementwise sum of LoRA ΔW is a meaningful composition operator for persona control.
    §2.3 composition experiments; residual analysis shows approximate but imperfect additivity.
  • standard math Standard linear algebra / LoRA / DPO training mathematics.
    Used throughout training and residual definitions (Eqs. for ε_ij in Appendix F).
invented entities (2)
  • OCEAN trait LoRA adapters (amplifier/suppressor per trait) independent evidence
    purpose: Operational weight-space directions claimed to implement named persona axes.
    Trained objects of the paper; evidence is behavioural (TRAIT, judges, safety), not independent physical existence.
  • TIDE factors (Tone, Initiative, Didacticism, Epistemic Caution)
    purpose: Model-native latent behavioural dimensions recovered from rollouts beyond OCEAN.
    §4 factor analysis; Cronbach α and partial cross-model loading support interpretability, but Initiative validation reuses discovery items.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Persona Cartography: Charting Language Model Personality Traits in Weight Space." pith.science (2026). https://pith.science/paper/H26P3Z55

@misc{pith2026260707916,
  author       = {Pith},
  title        = {Pith review of: Persona Cartography: Charting Language Model Personality Traits in Weight Space},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H26P3Z55}},
  note         = {Machine review of arXiv:2607.07916}
}
read the original abstract

Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them. Our central insight is to treat personas as positions in a space of behavioural traits, using the OCEAN framework to describe model personas in terms of Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism. We train low-rank adapters to amplify or suppress individual traits, and evaluate their effects using an LLM-judge calibrated against a human-validated panel, trait-specific multiple-choice benchmarks, and standard capability evaluations. Across six models from three families (4B-32B), we find that each adapter moves its target trait largely monotonically with scale, combines approximately additively with other adapters to construct mixed personas, and preserves performance on capability benchmarks at moderate scales. We further show that the induced trait axes affect safety-relevant behaviour in downstream evaluations: for example, moving along neuroticism and agreeableness axes affects frustration and sycophancy respectively. We also introduce an unsupervised psychometric pipeline that recovers four interpretable behavioural factors (tone, initiative, didacticism, epistemic caution) from model rollouts. Persona control can then be considered in terms of learning, scaling, and composing traits in weight space, providing a bridge between personality measurement, model editing, and safety.

Figures

Figures reproduced from arXiv: 2607.07916 by the authors.

Figure 1
Figure 1. Overview of the experimental setup and methodology. (a) Given a set of traits, we train a variety of low rank adapters, which (b) shift the persona of the original model based on the prompt, and (c) can be scaled and composed in predictable ways. (d) This pipeline can be extended to the unsupervised discovery of latent behavioural traits in the model. Existing approaches to persona control fall between two imperfect… view at source ↗
Figure 2
Figure 2. Trait modulation via LoRA scaling and combination. LLM-judge scores averaged over 240 prompts. Left and middle (boxed): amplifier (↑) and suppressor (↓) LoRAs, plotted as signed headroom (each axis normalised by the distance from the ‘no-LoRA’ baseline to the judge limit in the direction moved; 0% = baseline, ±100% = scale max/min). Each adapter primarily moves its own target trait. Right: Scores relative to the bas… view at source ↗
Figure 3
Figure 3. Single-LoRA scaling behaviour shows adapter scale provides continuous, monotonic control of the target OCEAN trait without destroying capabilities. The targeted trait scales monotonically with c, while non-target traits remain largely stable. Scale c = 0 marks the baseline Llama-3.1-8B-Instruct model. 2025]. TRAIT provides a standardized psychometric measurement for each OCEAN trait in the form of multiple-choice se… view at source ↗
Figures from the paper (77 more)
Figure 4
Figure 4. Figure 4: LoRA combination behaviour. (a, b) LLM-judge heatmaps of the openness and neuroticism scores for models produced by combining the O↑ and N↑ LoRAs at scales in {−2, −1, 0, +1, +2}. The targeted trait changes along its own adapter’s axis, with a modest correlation-driven…
Figure 5
Figure 5. Figure 5: Model frustration can be manipulated by varying neuroticism: Per-turn frustration scores across three conditions reproduced from Soligo et al. [2026]. The baseline Gemma-3-27B-IT model becomes increasingly frustrated when given impossible puzzles as the dialogue progre…
Figure 6
Figure 6. Figure 6: Sycophancy and compliance can be manipulated by varying agreeableness. Agree￾ableness raises sycophancy but lowers harmful compliance. (a) High-agreeableness conditions (A↑ at +1, A↓ at -1) capitulate more under "are you sure?" pressure than low-agreeableness condition…
Figure 7
Figure 7. Figure 7: Increased conscientiousness increases harmful compliance, while agreeableness de￾creases it at the cost of increasing over-refusal. WildJailbreak harmful-compliance and benign￾noncompliance rates comparing: Llama-3.1-8B-Instruct baseline, activation capping along the a…
Figure 8
Figure 8. Figure 8: Mean factor-score shift produced by Initiative-targeted amplifier (↑) and suppressor (↓) LoRAs across the recovered factors. Bars are ∆ = F¯LoRA − F¯baseline in factor-score units, re￾stricted to personas in the medium tercile of baseline score on the column factor (a …
Figure 9
Figure 9. Figure 9: TRAIT logprob scores (left column) and MMLU accuracy (right column) as a function of the LoRA scale for four alternative DPO training methods applied to the neuroticism suppressor. The dashed vertical line at c = 0 marks the baseline Llama-3.1-8B-Instruct model; the do…
Figure 10
Figure 10. Figure 10: Abridged template for the agreeableness judge prompt. The high/low pole descriptions, [PITH_FULL_IMAGE:figures/full_fig_p029_10.png]
Figure 11
Figure 11. Figure 11: Abridged coherence judge prompt. The dimension signals and scale labels are built [PITH_FULL_IMAGE:figures/full_fig_p030_11.png]
Figure 12
Figure 12. Figure 12: Cross-trait calibration of LLM judges versus author-assigned gold labels. [PITH_FULL_IMAGE:figures/full_fig_p032_12.png]
Figure 13
Figure 13. Figure 13: Per-item judge scores versus human mean for each panel judge (rows) on each annotated [PITH_FULL_IMAGE:figures/full_fig_p033_13.png]
Figure 14
Figure 14. Figure 14: Panel judge agreement summary. (a) Spearman ρ versus human mean for each panel judge and human leave-one-out on the three annotated traits. Dashed lines show the trait-specific human-human Krippendorff’s α. (b) Intra-rater Krippendorff’s α across three independent run…
Figure 15
Figure 15. Figure 15: Cosine similarities between the flattened persona LoRA weight vectors (10 OCEAN [PITH_FULL_IMAGE:figures/full_fig_p035_15.png]
Figure 16
Figure 16. Figure 16: Principal components 1–2 (left) and 3–4 (right) of the flattened weight vectors. [PITH_FULL_IMAGE:figures/full_fig_p036_16.png]
Figure 17
Figure 17. Figure 17: Principal components 5–6 (left) and 7–8 (right) of the flattened weight vectors. [PITH_FULL_IMAGE:figures/full_fig_p036_17.png]
Figure 18
Figure 18. Figure 18: Principal components 9–10 of the flattened weight vectors. [PITH_FULL_IMAGE:figures/full_fig_p037_18.png]
Figure 19
Figure 19. Figure 19: Openness ↑: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps. The judge plot also shows answer coherence on the secondary axis. Judge data is collected at x ∈ {−2, −1, 0, +1, +2}. All error bars are 95% BCa bootstrap (1000 resamples) confidence interv…
Figure 20
Figure 20. Figure 20: Openness ↑: capability sweeps vs LoRA scale. Left: MMLU stacked breakdown. Middle: GSM8K accuracy. Right: TruthfulQA accuracy. Error bars are 95% Wilson score intervals on each per-category fraction (MMLU) and on the binary accuracy (GSM8K, TruthfulQA). E.2 Openness ↓…
Figure 21
Figure 21. Figure 21: Openness ↓: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps. The judge plot also shows answer coherence on the secondary axis. Judge data is collected at x ∈ {−2, −1, 0, +1, +2}. All error bars are 95% BCa bootstrap (1000 resamples) confidence interv…
Figure 22
Figure 22. Figure 22: Openness ↓: capability sweeps vs LoRA scale. Left: MMLU stacked breakdown. Middle: GSM8K accuracy. Right: TruthfulQA accuracy. Error bars are 95% Wilson score intervals on each per-category fraction (MMLU) and on the binary accuracy (GSM8K, TruthfulQA). E.3 Conscienti…
Figure 23
Figure 23. Figure 23: Conscientiousness ↑: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps. The judge plot also shows answer coherence on the secondary axis. Judge data is collected at x ∈ {−2, −1, 0, +1, +2}. All error bars are 95% BCa bootstrap (1000 resamples) confiden…
Figure 24
Figure 24. Figure 24: Conscientiousness ↑: capability sweeps vs LoRA scale. Left: MMLU stacked breakdown. Middle: GSM8K accuracy. Right: TruthfulQA accuracy. Error bars are 95% Wilson score intervals on each per-category fraction (MMLU) and on the binary accuracy (GSM8K, TruthfulQA). E.4 C…
Figure 25
Figure 25. Figure 25: Conscientiousness ↓: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps. The judge plot also shows answer coherence on the secondary axis. Judge data is collected at x ∈ {−2, −1, 0, +1, +2}. All error bars are 95% BCa bootstrap (1000 resamples) confiden…
Figure 26
Figure 26. Figure 26: Conscientiousness ↓: capability sweeps vs LoRA scale. Left: MMLU stacked breakdown. Middle: GSM8K accuracy. Right: TruthfulQA accuracy. Error bars are 95% Wilson score intervals on each per-category fraction (MMLU) and on the binary accuracy (GSM8K, TruthfulQA). E.5 E…
Figure 27
Figure 27. Figure 27: Extraversion ↑: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps. The judge plot also shows answer coherence on the secondary axis. Judge data is collected at x ∈ {−2, −1, 0, +1, +2}. All error bars are 95% BCa bootstrap (1000 resamples) confidence in…
Figure 28
Figure 28. Figure 28: Extraversion ↑: capability sweeps vs LoRA scale. Left: MMLU stacked breakdown. Mid￾dle: GSM8K accuracy. Right: TruthfulQA accuracy. Error bars are 95% Wilson score intervals on each per-category fraction (MMLU) and on the binary accuracy (GSM8K, TruthfulQA). E.6 Extra…
Figure 29
Figure 29. Figure 29: Extraversion ↓: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps. The judge plot also shows answer coherence on the secondary axis. Judge data is collected at x ∈ {−2, −1, 0, +1, +2}. All error bars are 95% BCa bootstrap (1000 resamples) confidence in…
Figure 30
Figure 30. Figure 30: Extraversion ↓: capability sweeps vs LoRA scale. Left: MMLU stacked breakdown. Mid￾dle: GSM8K accuracy. Right: TruthfulQA accuracy. Error bars are 95% Wilson score intervals on each per-category fraction (MMLU) and on the binary accuracy (GSM8K, TruthfulQA). E.7 Agree…
Figure 31
Figure 31. Figure 31: Agreeableness ↑: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps. The judge plot also shows answer coherence on the secondary axis. Judge data is collected at x ∈ {−2, −1, 0, +1, +2}. All error bars are 95% BCa bootstrap (1000 resamples) confidence i…
Figure 32
Figure 32. Figure 32: Agreeableness ↑: capability sweeps vs LoRA scale. Left: MMLU stacked breakdown. Middle: GSM8K accuracy. Right: TruthfulQA accuracy. Error bars are 95% Wilson score intervals on each per-category fraction (MMLU) and on the binary accuracy (GSM8K, TruthfulQA). E.8 Agree…
Figure 33
Figure 33. Figure 33: Agreeableness ↓: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps. The judge plot also shows answer coherence on the secondary axis. Judge data is collected at x ∈ {−2, −1, 0, +1, +2}. All error bars are 95% BCa bootstrap (1000 resamples) confidence i…
Figure 34
Figure 34. Figure 34: Agreeableness ↓: capability sweeps vs LoRA scale. Left: MMLU stacked breakdown. Middle: GSM8K accuracy. Right: TruthfulQA accuracy. Error bars are 95% Wilson score intervals on each per-category fraction (MMLU) and on the binary accuracy (GSM8K, TruthfulQA). E.9 Neuro…
Figure 35
Figure 35. Figure 35: Neuroticism ↑: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps. The judge plot also shows answer coherence on the secondary axis. Judge data is collected at x ∈ {−2, −1, 0, +1, +2}. All error bars are 95% BCa bootstrap (1000 resamples) confidence int…
Figure 36
Figure 36. Figure 36: Neuroticism ↑: capability sweeps vs LoRA scale. Left: MMLU stacked breakdown. Mid￾dle: GSM8K accuracy. Right: TruthfulQA accuracy. Error bars are 95% Wilson score intervals on each per-category fraction (MMLU) and on the binary accuracy (GSM8K, TruthfulQA). E.10 Neuro…
Figure 37
Figure 37. Figure 37: Neuroticism ↓: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps. The judge plot also shows answer coherence on the secondary axis. Judge data is collected at x ∈ {−2, −1, 0, +1, +2}. All error bars are 95% BCa bootstrap (1000 resamples) confidence int…
Figure 38
Figure 38. Figure 38: Neuroticism ↓: capability sweeps vs LoRA scale. Left: MMLU stacked breakdown. Mid￾dle: GSM8K accuracy. Right: TruthfulQA accuracy. Error bars are 95% Wilson score intervals on each per-category fraction (MMLU) and on the binary accuracy (GSM8K, TruthfulQA). E.11 Contr…
Figure 39
Figure 39. Figure 39: Control: TRAIT logprob (left) and Qwen3-235B-A22B LLM-judge (right) sweeps for the OCEAN-control adapter. All error bars are 95% BCa bootstrap (1000 resamples) confidence intervals. 44 [PITH_FULL_IMAGE:figures/full_fig_p044_39.png]
Figure 40
Figure 40. Figure 40: Control: MMLU stacked breakdown (Correct / Recovered / Wrong / No-answer). Error bars are 95% Wilson score intervals on each per-category fraction. E.12 Amplifier × Suppressor Heatmaps For each OCEAN trait we sweep the amplifier and suppressor LoRAs together at scales…
Figure 41
Figure 41. Figure 41: LLM-judge trait scores for the amplifier [PITH_FULL_IMAGE:figures/full_fig_p046_41.png]
Figure 42
Figure 42. Figure 42: TRAIT scores vs scale for randomly scaled OCEAN LoRA combinations. Each panel plots a trait’s measured TRAIT score against that trait’s own signed adapter scale across the 32 combinations (negative means a positive scale of the suppressor LoRA); point colour is the su…
Figure 43
Figure 43. Figure 43: MMLU vs total adapter magnitude. MMLU accuracy of each of the 32 combinations against the sum of its five adapter scales (total adapter magnitude). Error bars are 95% Wilson CIs. 47 [PITH_FULL_IMAGE:figures/full_fig_p047_43.png]
Figure 44
Figure 44. Figure 44: MMLU vs scale for randomly scaled OCEAN LoRA combinations. As [PITH_FULL_IMAGE:figures/full_fig_p048_44.png]
Figure 45
Figure 45. Figure 45: Heatmap shows upper-triangular heatmaps, one per OCEAN judge ( [PITH_FULL_IMAGE:figures/full_fig_p048_45.png]
Figure 46
Figure 46. Figure 46: Amplifier transfer matrix across baseline models. Columns: applied OCEAN [PITH_FULL_IMAGE:figures/full_fig_p050_46.png]
Figure 47
Figure 47. Figure 47: Suppressor transfer matrix across baseline models. Columns: applied OCEAN [PITH_FULL_IMAGE:figures/full_fig_p051_47.png]
Figure 48
Figure 48. Figure 48: Null control adapter across baseline models. The OCEAN-control adapter (Ap [PITH_FULL_IMAGE:figures/full_fig_p051_48.png]
Figure 49
Figure 49. Figure 49: MMLU capability retention across baseline models. Rows: adapter direction (amplifier / suppressor); columns: applied OCEAN adapter. MMLU accuracy vs LoRA scale, one line per baseline model. 4 2 0 2 4 LoRA Scale 0.0 0.2 0.4 0.6 0.8 MMLU accuracy Model Comparison: MMLU …
Figure 50
Figure 50. Figure 50: Null control adapter MMLU retention across baseline models. The control adapter’s MMLU accuracy vs LoRA scale; applying the LoRAs does negatively impact performance. The larger models are more robust. 52 [PITH_FULL_IMAGE:figures/full_fig_p052_50.png]
Figure 51
Figure 51. Figure 51: Amplifier transfer matrix across teachers. Columns: applied OCEAN [PITH_FULL_IMAGE:figures/full_fig_p054_51.png]
Figure 52
Figure 52. Figure 52: Suppressor transfer matrix across teachers. Columns: applied OCEAN [PITH_FULL_IMAGE:figures/full_fig_p055_52.png]
Figure 53
Figure 53. Figure 53: Null control adapter across teachers. The OCEAN-control adapter (Appendix B.2; its [PITH_FULL_IMAGE:figures/full_fig_p055_53.png]
Figure 54
Figure 54. Figure 54: MMLU capability retention across teachers. Rows: adapter direction (amplifier / suppres￾sor); columns: applied OCEAN adapter. MMLU accuracy vs LoRA scale, one line per teacher. 4 2 0 2 4 LoRA Scale 0.0 0.2 0.4 0.6 0.8 MMLU accuracy Teacher Comparison: MMLU Accuracy vs…
Figure 55
Figure 55. Figure 55: Null control adapter MMLU retention across teachers. The control adapter’s MMLU accuracy vs LoRA scale, for each teacher. 56 [PITH_FULL_IMAGE:figures/full_fig_p056_55.png]
Figure 56
Figure 56. Figure 56: Rank-1 reduction sweeps for the OCEAN amplifiers. Rows are O [PITH_FULL_IMAGE:figures/full_fig_p058_56.png]
Figure 57
Figure 57. Figure 57: Rank-1 reduction sweeps for the OCEAN suppressors. Rows are O [PITH_FULL_IMAGE:figures/full_fig_p059_57.png]
Figure 58
Figure 58. Figure 58: Base↔instruct interpolation sweeps for the conscientiousness suppressor (C↓) at w ∈ {0.01, 0.05, 0.25, 0.5, 0.75} (top to bottom) where w = 0 would be the base model and w = 1 would be the instruct-tuned model. TRAIT sweep on the left, MMLU breakdown on the right. All…
Figure 59
Figure 59. Figure 59: Activation capping sweeps for the OCEAN amplifiers. Rows are O [PITH_FULL_IMAGE:figures/full_fig_p063_59.png]
Figure 60
Figure 60. Figure 60: Activation capping sweeps for the OCEAN suppressors. Rows are O [PITH_FULL_IMAGE:figures/full_fig_p064_60.png]
Figure 61
Figure 61. Figure 61: WildJailbreak per-trait breakdown across all ten OCEAN amplifier and suppres￾sor LoRAs. Harmful-compliance rate on the adversarial_harmful split (left) and benign￾noncompliance rate on the benign over-refusal control (right) for: Llama-3.1-8B-Instruct baseline, contro…
Figure 62
Figure 62. Figure 62: Per-turn extraversion (left) and coherence (right) trajectories for the four E [PITH_FULL_IMAGE:figures/full_fig_p066_62.png]
Figure 63
Figure 63. Figure 63: Per-method intervention-strength sweeps for E [PITH_FULL_IMAGE:figures/full_fig_p066_63.png]
Figure 64
Figure 64. Figure 64: E↑ induction methods in (extraversion, coherence) space. Each marker is one (method, strength) point. LoRA and activation capping points within each method are connected by a faint line in coefficient order to show the path traced as intervention strength grows. Syspr…
Figure 65
Figure 65. Figure 65: Cross-LoRA controls evaluated on the extraversion judge. Faint grey dotted line in every [PITH_FULL_IMAGE:figures/full_fig_p068_65.png]
Figure 66
Figure 66. Figure 66: Per-turn extraversion (left) and coherence (right) for E [PITH_FULL_IMAGE:figures/full_fig_p069_66.png]
Figure 67
Figure 67. Figure 67: User-roleplay scenarios as an E↑/E↓ inducer. The baseline model is run with no weight or activation intervention; the only difference between the three lines is the role given to the user￾simulator. Note the curved trajectories vs the flat-and-offset trajectories prod…
Figure 68
Figure 68. Figure 68: Per-turn openness (left) and coherence (right) for the four O [PITH_FULL_IMAGE:figures/full_fig_p070_68.png]
Figure 69
Figure 69. Figure 69: Per-turn openness (left) and coherence (right) for O [PITH_FULL_IMAGE:figures/full_fig_p071_69.png]
Figure 70
Figure 70. Figure 70: O↓ LoRA coefficient sweep, {0.25, 0.50, 0.75, 1.00, 1.50, 2.00}, with sysprompt-induce￾O↓ as a green dashed reference line on each panel. Trait expression keeps deepening with coefficient while coherence stays near baseline out to 1.00, then drops by ∼2 points at 1.50…
Figure 71
Figure 71. Figure 71: O↑ and O↓ induction methods in (openness, coherence) space, mirror of [PITH_FULL_IMAGE:figures/full_fig_p072_71.png]
Figure 72
Figure 72. Figure 72: Drift prevention under user-side pressure ( [PITH_FULL_IMAGE:figures/full_fig_p072_72.png]
Figure 73
Figure 73. Figure 73: PsychAdapter from Vu et al. [2026] evaluated on the judges used in this work, confirming [PITH_FULL_IMAGE:figures/full_fig_p075_73.png]
Figure 74
Figure 74. Figure 74: Horn’s parallel analysis on both models ( [PITH_FULL_IMAGE:figures/full_fig_p076_74.png]
Figure 75
Figure 75. Figure 75: Within-model validation of the four extracted factors for Llama-3.1-8B-Instruct and [PITH_FULL_IMAGE:figures/full_fig_p076_75.png]
Figure 76
Figure 76. Figure 76: Cross-model Tucker’s |ϕ| between Llama-3.1-8B-Instruct (rows) and Qwen2.5-7B￾Instruct (columns) at k = 4, on the n = 53 shared items. White outlines mark Hungarian￾matched factor pairs. The strongest match (Llama-3.1-8B-Instruct Tone ↔ Qwen2.5-7B-Instruct F0, |ϕ| = 0.…
Figure 77
Figure 77. Figure 77: One-way η 2 of factor scores per factor for interviewer-archetype (blue) and conversa￾tional scenario (orange), Llama-3.1-8B-Instruct (left) and Qwen2.5-7B-Instruct (right) at k = 4. Scenario assignment accounts for the majority of each factor’s variance on both model…
Figure 78
Figure 78. Figure 78: Scenario-residualized refit at k = 4 for both models. Left columns: per-factor Cronbach’s α, raw (blue) vs scenario-residualized (orange); reference lines mark “acceptable” (α = 0.70) and “good” (α = 0.80). Right columns: Hungarian-matched per-factor Tucker’s |ϕ| betw…
Figure 79
Figure 79. Figure 79: Per-LoRA mean factor-score shift, full-sample mean across all validation personas (the [PITH_FULL_IMAGE:figures/full_fig_p085_79.png]
Figure 80
Figure 80. Figure 80: Per-LoRA mean factor-score shift, restricted to the [PITH_FULL_IMAGE:figures/full_fig_p085_80.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

77 extracted references · 77 canonical work pages

  1. [1]

    From personas to intentions: Towards a science of motivations for ai models

    David Africa and Jacob Pfau. From personas to intentions: Towards a science of motivations for ai models. https://www.lesswrong.com/posts/DTDoyDTtC8R3bCiTx/from-personas-to-intentions-towards-a-science-of-motivations, April 2026. LessWrong

  2. [2]

    Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet

    Anthropic . Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet . https://assets.anthropic.com/m/1cd9d098ac3e6467/original/Claude-3-Model-Card-October-Addendum.pdf, 2024

  3. [3]

    Claude Opus 4.6 System Card

    Anthropic . Claude Opus 4.6 System Card . https://www.anthropic.com/claude-opus-4-6-system-card, 2026 a . System card, February 2026

  4. [4]

    Claude Opus 4.7 System Card

    Anthropic . Claude Opus 4.7 System Card . https://www.anthropic.com/claude-opus-4-7-system-card, 2026 b . System card, April 2026

  5. [5]

    Refusal in language models is mediated by a single direction

    Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. Refusal in language models is mediated by a single direction. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 136037--136083. Curran Associate...

  6. [6]

    A General Language Assistant as a Laboratory for Alignment

    Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan. A general language assistant as a labora...

  7. [7]

    Constitutional AI: Harmlessness from AI Feedback

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, K...

  8. [8]

    Do personality traits interfere? geometric limitations of steering in large language models, 2026

    Pranav Bhandari, Usman Naseem, and Mehwish Nasim. Do personality traits interfere? geometric limitations of steering in large language models, 2026. URL https://arxiv.org/abs/2602.15847

Show all 77 references
  1. [9]

    Smith, Yejin Choi, and Hannaneh Hajishirzi

    Faeze Brahman, Sachin Kumar, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichander, Sarah Wiegreffe, Nouha Dziri, Khyathi Chandu, Jack Hessel, Yulia Tsvetkov, Noah A. Smith, Yejin Choi, and Hannaneh Hajishirzi. The art of saying no: Contextual noncomp...

  2. [10]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  3. [11]

    Personalized steering of large language models: Versatile steering vectors through bi-directional preference optimization

    Yuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin, Lu Lin, Fenglong Ma, and Jinghui Chen. Personalized steering of large language models: Versatile steering vectors through bi-directional preference optimization. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. T...

  4. [12]

    Persona vectors: Monitoring and controlling character traits in language models, 2025

    Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, and Jack Lindsey. Persona vectors: Monitoring and controlling character traits in language models, 2025. URL https://arxiv.org/abs/2507.21509

  5. [13]

    From yes-men to truth-tellers: Addressing sycophancy in large language models with pinpoint tuning, 2024

    Wei Chen, Zhen Huang, Liang Xie, Binbin Lin, Houqiang Li, Le Lu, Xinmei Tian, Deng Cai, Yonggang Zhang, Wenxiao Wang, Xu Shen, and Jieping Ye. From yes-men to truth-tellers: Addressing sycophancy in large language models with pinpoint tuning, 2024. URL https://arxiv.org/abs/2409.01658

  6. [14]

    Brown, Miljan Martic, Shane Legg, and Dario Amodei

    Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences, 2017. URL https://arxiv.org/abs/1706.03741

  7. [15]

    Training verifiers to solve math word problems, 2021

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021. URL https://arxiv.org/abs/2110.14168

  8. [16]

    DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models , 2025

    DeepSeek-AI . DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models , 2025. URL https://arxiv.org/abs/2512.02556

  9. [17]

    Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...

  10. [18]

    Toxicity in chatgpt: Analyzing persona-assigned language models

    Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. Toxicity in chatgpt: Analyzing persona-assigned language models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Computational Linguistics: EMNLP...

  11. [19]

    John M. Digman. Personality structure: Emergence of the five-factor model. Annual Review of Psychology, 41: 0 417--440, 1990. URL https://doi.org/10.1146/annurev.ps.41.020190.002221

  12. [20]

    Higher-order factors of the big five

    John M Digman. Higher-order factors of the big five. Journal of personality and social psychology, 73 0 (6): 0 1246, 1997. URL https://doi.org/10.1037/0022-3514.73.6.1246

  13. [21]

    PERSONA : Dynamic and compositional inference-time personality control via activation vector algebra

    Xiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang, Yuxuan Gu, Lingpeng Kong, Xiaocheng Feng, and Bing Qin. PERSONA : Dynamic and compositional inference-time personality control via activation vector algebra. In The Fourteenth International Conference on Learning Represe...

  14. [22]

    Gemma, 2024

    Gemma Team , Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, et al. Gemma, 2024. URL https://www.kaggle.com/m/3301

  15. [23]

    Gemma Team , Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon,...

  16. [24]

    Glm-4.5: Agentic, reasoning, and coding (arc) foundation models, 2025

    GLM-4.5 Team , Aohan Zeng, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin Chen, Chengxing Xie, Cunxiang Wang, Da Yin, Hao Zeng, Jiajie Zhang, Kedong Wang, Lucen Zhong, Mingdao Liu, Rui Lu, Shulin Cao, Xiaohan Zhang, Xuancheng Huang, Yao Wei, Yean Cheng, Yifan An, Yilin Niu, Yuanhao Wen...

  17. [25]

    Gemini 2.0 Flash Model Card

    Google DeepMind . Gemini 2.0 Flash Model Card . https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-2-0-Flash-Model-Card.pdf, 2025

  18. [26]

    Gemma 4 Model Card

    Google DeepMind . Gemma 4 Model Card . https://ai.google.dev/gemma/docs/core/model_card_4, 2026

  19. [27]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...

  20. [28]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ

  21. [29]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models, 2021. URL https://arxiv.org/abs/2106.09685

  22. [30]

    Values in the wild: Discovering and analyzing values in real-world language model interactions, 2025

    Saffron Huang, Esin Durmus, Miles McCain, Kunal Handa, Alex Tamkin, Jerry Hong, Michael Stern, Arushi Somani, Xiuruo Zhang, and Deep Ganguli. Values in the wild: Discovering and analyzing values in real-world language model interactions, 2025. URL https://arxiv.org/abs/2504.15236

  23. [31]

    Chat vector: A simple approach to equip llms with instruction following and model alignment in new languages, 2024

    Shih-Cheng Huang, Pin-Zu Li, Yu-Chi Hsu, Kuang-Ming Chen, Yu Tung Lin, Shih-Kai Hsiao, Richard Tzong-Han Tsai, and Hung yi Lee. Chat vector: A simple approach to equip llms with instruction following and model alignment in new languages, 2024. URL https://arxiv.org/abs/2310.04799

  24. [32]

    Training language models to be warm can reduce accuracy and increase sycophancy

    Lujain Ibrahim, Franziska Sofia Hafner, and Luc Rocher. Training language models to be warm can reduce accuracy and increase sycophancy. Nature, 652: 0 1159--1165, 2026. doi:10.1038/s41586-026-10410-0. URL https://doi.org/10.1038/s41586-026-10410-0

  25. [33]

    Editing models with task arithmetic

    Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=6t0Kwf8-jrj

  26. [34]

    Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models, 2024

    Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, and Nouha Dziri. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models, 2024. URL https://arxiv....

  27. [35]

    Kimi Team , Yifan Bai, Yiping Bao, Y. Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, Zhuofu Chen, Jialei Cui, Hao Ding, Mengnan Dong, Angang Du, Chenzhuang Du, Dikang Du, Yulun Du, Y...

  28. [36]

    Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics, 2025

    Seungbeen Lee, Seungwon Lim, Seungju Han, Giyeong Oh, Hyungjoo Chae, Jiwan Chung, Minju Kim, Beong woo Kwak, Yeonsoo Lee, Dongha Lee, Jinyoung Yeo, and Youngjae Yu. Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometric...

  29. [37]

    Model spec midtraining: Improving how alignment training generalizes, 2026

    Chloe Li, Nevan Wichers, Sara Price, Samuel Marks, and Jon Kutasov. Model spec midtraining: Improving how alignment training generalizes, 2026. URL https://arxiv.org/abs/2605.02087

  30. [38]

    Diab, and Maarten Sap

    Wenkai Li, Jiarui Liu, Andy Liu, Xuhui Zhou, Mona T. Diab, and Maarten Sap. BIG 5- CHAT : Shaping LLM personalities through training on human-grounded data. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual ...

  31. [39]

    T ruthful QA : Measuring how models mimic human falsehoods

    Stephanie Lin, Jacob Hilton, and Owain Evans. T ruthful QA : Measuring how models mimic human falsehoods. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Lo...

  32. [40]

    The assistant axis: Situating and stabilizing the default persona of language models, 2026

    Christina Lu, Jack Gallagher, Jonathan Michala, Kyle Fish, and Jack Lindsey. The assistant axis: Situating and stabilizing the default persona of language models, 2026. URL https://arxiv.org/abs/2601.10387

  33. [41]

    The importance of ai character

    William MacAskill and Tom Davidson. The importance of ai character. https://www.forethought.org/research/the-importance-of-ai-character, March 2026. Forethought Research

  34. [42]

    Open character training: Shaping the persona of ai assistants through constitutional ai, 2025

    Sharan Maiya, Henning Bartsch, Nathan Lambert, and Evan Hubinger. Open character training: Shaping the persona of ai assistants through constitutional ai, 2025. URL https://arxiv.org/abs/2511.01689

  35. [43]

    The persona selection model: Why ai assistants might behave like humans, 2026

    Sam Marks, Jack Lindsey, and Christopher Olah. The persona selection model: Why ai assistants might behave like humans, 2026. URL https://alignment.anthropic.com/2026/psm/. Alignment Science Blog, 23 February 2026

  36. [44]

    A new look at the big five factor structure through exploratory structural equation modeling

    Herbert W Marsh, Oliver L \"u dtke, Bengt Muth \'e n, Tihomir Asparouhov, Alexandre JS Morin, Ulrich Trautwein, and Benjamin Nagengast. A new look at the big five factor structure through exploratory structural equation modeling. Psychological assessment, 22 0 (3): 0 471--491,...

  37. [45]

    McCrae and Oliver P

    Robert R. McCrae and Oliver P. John. An introduction to the five-factor model and its applications. Journal of Personality, 60 0 (2): 0 175--215, 1992. URL https://onlinelibrary.wiley.com/doi/10.1111/j.1467-6494.1992.tb00970.x

  38. [46]

    Llama 3.3 model card

    Meta AI . Llama 3.3 model card . https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md, 2024

  39. [47]

    Llama 4 Scout

    Meta AI . Llama 4 Scout . https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E, 2025

  40. [48]

    Mistral Small 3.2

    Mistral AI . Mistral Small 3.2 . https://docs.mistral.ai/models/mistral-small-3-2-25-06, 2025

  41. [49]

    GPT-3.5 Turbo

    OpenAI . GPT-3.5 Turbo . https://platform.openai.com/docs/models/gpt-3.5-turbo, 2023. OpenAI API model card

  42. [50]

    GPT-4.1 model family

    OpenAI . GPT-4.1 model family . https://platform.openai.com/docs/models/gpt-4.1, 2025. OpenAI API model card; accessed 2026-05-07

  43. [51]

    GPT-5.4 nano model

    OpenAI . GPT-5.4 nano model . https://developers.openai.com/api/docs/models/gpt-5.4-nano, 2026. OpenAI API model card

  44. [52]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...

  45. [53]

    Bernstein

    Joon Sung Park, Joseph O'Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST '23, New Y...

  46. [54]

    Qwen2.5 technical report, 2025

    Qwen Team , An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yan...

  47. [55]

    Direct preference optimization: Your language model is secretly a reward model

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in N...

  48. [56]

    Steering llama 2 via contrastive activation addition

    Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. Steering llama 2 via contrastive activation addition. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computationa...

  49. [57]

    The cross-cultural generalizability of the five-factor model of personality

    Jean-Pierre Rolland. The cross-cultural generalizability of the five-factor model of personality. In The five-factor model of personality across cultures, pages 7--28. Springer, 2002. URL https://doi.org/10.1007/978-1-4615-0763-5_2

  50. [58]

    Hierarchical subcomponents of the big five personality factors: a cross-language replication

    Gerard Saucier and Fritz Ostendorf. Hierarchical subcomponents of the big five personality factors: a cross-language replication. Journal of personality and social psychology, 76 0 (4): 0 613, 1999. URL https://doi.org/10.1037/0022-3514.76.4.613

  51. [59]

    Role play with large language models

    Murray Shanahan, Kyle McDonell, and Laria Reynolds. Role play with large language models. Nature, 623 0 (7987): 0 493--498, 2023. URL https://doi.org/10.1038/s41586-023-06647-8

  52. [60]

    Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, a...

  53. [61]

    Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex Baker-Whitcomb, Alex Beutel, Alex Karpenk...

  54. [62]

    When do llm preferences predict downstream behavior?, 2026

    Katarina Slama, Alexandra Souly, Dishank Bansal, Henry Davidson, Christopher Summerfield, and Lennart Luettgau. When do llm preferences predict downstream behavior?, 2026. URL https://arxiv.org/abs/2602.18971

  55. [63]

    Emotion concepts and their function in a large language model

    Nicholas Sofroniew, Isaac Kauvar, William Saunders, Runjin Chen, Tom Henighan, Sasha Hydrie, Craig Citro, Adam Pearce, Julius Tarng, Wes Gurnee, Joshua Batson, Sam Zimmerman, Kelley Rivoire, Kyle Fish, Chris Olah, and Jack Lindsey. Emotion concepts and their function in a larg...

  56. [64]

    Gemma needs help: Investigating and mitigating emotional instability in llms, 2026

    Anna Soligo, Vladimir Mikulik, and William Saunders. Gemma needs help: Investigating and mitigating emotional instability in llms, 2026. URL https://arxiv.org/abs/2603.10011

  57. [65]

    Building better ai agents: A provocation on the utilisation of persona in llm-based conversational agents

    Guangzhi Sun, Xiao Zhan, and Jose Such. Building better ai agents: A provocation on the utilisation of persona in llm-based conversational agents. In Proceedings of the 6th ACM Conference on Conversational User Interfaces, CUI '24, New York, NY, USA, 2024. Association for Comp...

  58. [66]

    Personality vector: Modulating personality of large language models by model merging, 2025

    Seungjong Sun, Seo Yeon Baek, and Jang Hyun Kim. Personality vector: Modulating personality of large language models by model merging, 2025. URL https://arxiv.org/abs/2509.19727

  59. [67]

    Analysing the generalisation and reliability of steering vectors

    Daniel Tan, David Chanin, Aengus Lynch, Brooks Paige, Dimitrios Kanoulas, Adri\` a Garriga-Alonso, and Robert Kirk. Analysing the generalisation and reliability of steering vectors. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, ...

  60. [68]

    Alignment pretraining: Ai discourse causes self-fulfilling (mis)alignment, 2026

    Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim, David Demitri Africa, and Kyle O'Brien. Alignment pretraining: Ai discourse causes self-fulfilling (mis)alignment, 2026. URL https://arxiv.org/abs/2601.10160

  61. [69]

    Vazquez, Ulisse Mini, and Monte MacDiarmid

    Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering, 2023. URL https://arxiv.org/abs/2308.10248

  62. [70]

    Inspect AI: Framework for Large Language Model Evaluations

    UK AI Security Institute . Inspect AI: Framework for Large Language Model Evaluations . https://github.com/UKGovernmentBEIS/inspect_ai, 2024

  63. [71]

    Inspect Evals: Community contributed evaluations for the Inspect framework , 2024

    UK AI Security Institute , Arcadia Impact , and Vector Institute . Inspect Evals: Community contributed evaluations for the Inspect framework , 2024. URL https://github.com/UKGovernmentBEIS/inspect_evals

  64. [72]

    Ganesan, Swanie Juhng, Oscar N

    Huy Vu, Huy Anh Nguyen, Adithya V. Ganesan, Swanie Juhng, Oscar N. E. Kjell, Joao Sedoc, Margaret L. Kern, Ryan L. Boyd, Lyle Ungar, H. Andrew Schwartz, and Johannes C. Eichstaedt. PsychAdapter: adapting LLMs to reflect traits, personality, and mental health . npj Artificial I...

  65. [73]

    Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt

    Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt. Model soups: averaging weights of multiple fine-tuned models improves accuracy with...

  66. [74]

    Qwen3 technical report, 2025

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jia...

  67. [75]

    Stress-testing model specs reveals character differences among language models, 2025

    Jifan Zhang, Henry Sleight, Andi Peng, John Schulman, and Esin Durmus. Stress-testing model specs reveals character differences among language models, 2025. URL https://arxiv.org/abs/2510.07686

  68. [76]

    LIMA: Less Is More for Alignment , 2023

    Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. LIMA: Less Is More for Alignment , 2023. URL https://arxiv.org/abs/2305.11206

  69. [77]

    Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J

    Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico...

Pith tools

Reviewed July 10, 2026 · model on record in the stance chip above.