Pith. sign in

Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs

10 Pith papers cite this work. Polarity classification is still indexing.

10 Pith papers citing it
abstract

As AIs rapidly advance and become more agentic, the risk they pose is governed not only by their capabilities but increasingly by their propensities, including goals and values. Tracking the emergence of goals and values has proven a longstanding problem, and despite much interest over the years it remains unclear whether current AIs have meaningful values. We propose a solution to this problem, leveraging the framework of utility functions to study the internal coherence of AI preferences. Surprisingly, we find that independently-sampled preferences in current LLMs exhibit high degrees of structural coherence, and moreover that this emerges with scale. These findings suggest that value systems emerge in LLMs in a meaningful sense, a finding with broad implications. To study these emergent value systems, we propose utility engineering as a research agenda, comprising both the analysis and control of AI utilities. We uncover problematic and often shocking values in LLM assistants despite existing control measures. These include cases where AIs value themselves over humans and are anti-aligned with specific individuals. To constrain these emergent value systems, we propose methods of utility control. As a case study, we show how aligning utilities with a citizen assembly reduces political biases and generalizes to new scenarios. Whether we like it or not, value systems have already emerged in AIs, and much work remains to fully understand and control these emergent representations.

citation-role summary

background 1

citation-polarity summary

years

2026 9 2024 1

roles

background 1

polarities

support 1

representative citing papers

Artificial Persons

cs.CY · 2026-07-09 · conditional · novelty 8.0

Neither of Rawls' two moral powers requires sentience, so non-sentient AI systems can in principle be full political persons under the PCP.

Understanding Goal Generalisation in Sequential Reinforcement Learning

cs.LG · 2026-05-22 · unverdicted · novelty 6.0

Empirical analysis of over 100 sequential RL training pipelines across 250+ OOD environments finds salient features drive generalization and early goals persist, with latent policy gradients simulating latent variable evolution to predict OOD behavior from training history.

Probing Persona-Dependent Preferences in Language Models

cs.CL · 2026-05-13 · unverdicted · novelty 6.0

Linear probes on residual-stream activations identify a shared preference vector in LLMs that tracks choices across prompts and causally steers decisions even for anti-correlated personas.

Reducing Political Manipulation with Consistency Training

cs.CL · 2026-05-21 · unverdicted · novelty 5.0 · 2 refs

PCT is a reinforcement learning approach that trains LLMs for symmetric sentiment and helpfulness across paired opposing political prompts, reducing covert bias while preserving general performance.

citing papers explorer

Showing 10 of 10 citing papers.