Pith. sign in

REVIEW 2 cited by

ControlLM: Crafting Diverse Personalities for Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10151 v1 pith:WJ6XQOEJ submitted 2024-02-15 cs.CL

classification cs.CL
keywords languagemodelsbehaviorscontrollmmodelpersonalitycontroltraits
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As language models continue to scale in size and capability, they display an array of emerging behaviors, both beneficial and concerning. This heightens the need to control model behaviors. We hope to be able to control the personality traits of language models at the inference-time so as to have various character features, on top of which the requirements of different types of tasks can be met. Personality is a higher-level and more abstract behavioral representation for language models. We introduce ControlLM, which leverages differential activation patterns, derived from contrasting behavioral prompts in the model's latent space, to influence the model's personality traits at inference. This approach allows for the precise, real-time adjustment of model behavior. First, we demonstrate ControlLM's capacity to elicit diverse persona behaviors without any training, while precision control allows personality traits to closely match average human values. Subsequently, we showcase improved reasoning and question answering through selective amplification of beneficial attributes like conscientiousness and friendliness. We hope that this work will inspire research on controlling human-like behaviors of language models and provide insights for future research. Our code is publicly available at: https://github.com/wengsyx/ControlLM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

    cs.CL 2025-10 conditional novelty 6.0 of 10

    PsySET measures emotion and personality steering in LLMs across prompting, fine-tuning, and representation engineering, finding prompts most effective overall and emotion-specific safety trade-offs (e.g., joy weakens ...

  2. Understanding Persuasive Interactions between Generative Social Agents and Humans: The Knowledge-based Persuasion Model (KPM)

    cs.HC 2026-02 conditional novelty 4.0 of 10

    A proposed model says a generative social agent's self-, user-, and context-knowledge drives its persuasive behavior, which in turn shapes users' attitudes and behavior.

Pith tools