Pith. sign in

REVIEW 9 cited by

Moral Alignment for LLM Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.01639 v4 pith:WMMJYDC5 submitted 2024-10-02 cs.LG cs.AIcs.CY

classification cs.LGcs.AIcs.CY
keywords humanagentsmoralalignmentvaluesfine-tuningrewardsactivity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Decision-making agents based on pre-trained Large Language Models (LLMs) are increasingly being deployed across various domains of human activity. While their applications are currently rather specialized, several research efforts are underway to develop more generalist agents. As LLM-based systems become more agentic, their influence on human activity will grow and their transparency will decrease. Consequently, developing effective methods for aligning them to human values is vital. The prevailing practice in alignment often relies on human preference data (e.g., in RLHF or DPO), in which values are implicit, opaque and are essentially deduced from relative preferences over different model outputs. In this work, instead of relying on human feedback, we introduce the design of reward functions that explicitly and transparently encode core human values for Reinforcement Learning-based fine-tuning of foundation agent models. Specifically, we use intrinsic rewards for the moral alignment of LLM agents. We evaluate our approach using the traditional philosophical frameworks of Deontological Ethics and Utilitarianism, quantifying moral rewards for agents in terms of actions and consequences on the Iterated Prisoner's Dilemma (IPD) environment. We also show how moral fine-tuning can be deployed to enable an agent to unlearn a previously developed selfish strategy. Finally, we find that certain moral strategies learned on the IPD game generalize to several other matrix game environments. In summary, we demonstrate that fine-tuning with intrinsic rewards is a promising general solution for aligning LLM agents to human values, and it might represent a more transparent and cost-effective alternative to currently predominant alignment techniques.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Scalable Approach to Evaluating Moral Sensitivity in LLMs

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.

  2. Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate

    cs.AI 2025-10 conditional novelty 6.0 of 10

    In multi-agent debates over everyday moral dilemmas, GPT-4.1 almost never revises in simultaneous settings but conforms strongly in sequential settings, while Claude 3.7 and Gemini 2.0 Flash revise far more often.

  3. Localizing Persona Representations in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.

  4. When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Across morally framed prisoner's dilemmas and public goods games, none of nine LLMs consistently chooses the ethical action when it conflicts with payoff, with cooperation rates from 7.9% to 76.3%.

  5. Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Ten functional criteria, based on observable behavior rather than internal understanding, are offered for evaluating moral behavior in opaque large language model systems.

  6. Beyond Nash Equilibrium: Bounded Rationality of LLMs and humans in Strategic Decision-making

    cs.AI 2025-06 conditional novelty 5.0 of 10

    LLMs reproduce human heuristics like switching after a loss and cooperating when future rounds loom, but apply them more rigidly and adapt less than humans.

  7. Survival Games: Human-LLM Strategic Showdowns under Severe Resource Scarcity

    cs.HC 2025-05 reject novelty 5.0 of 10

    In a simulated survival game with two rule-based agents and one LLM-powered robot, DeepSeek models showed more detected unethical actions than OpenAI models, and jailbreak prompts sharply increased violations.

  8. Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey proposes a macro-meso-micro value framework for agentic AI alignment and maps applications, methods, and benchmarks onto it.

  9. A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A position paper arguing that LLM-based human-agent systems, not fully autonomous agents, should be the immediate goal for AI development.

Pith tools