Pith. sign in

REVIEW 1 cited by

Avoiding Wireheading with Value Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1605.03143 v1 pith:OTCIM2WH submitted 2016-05-10 cs.AI

classification cs.AI
keywords agentslearningreinforcementrewardwireheadingactionsagentconstraint
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How can we design good goals for arbitrarily intelligent agents? Reinforcement learning (RL) is a natural approach. Unfortunately, RL does not work well for generally intelligent agents, as RL agents are incentivised to shortcut the reward sensor for maximum reward -- the so-called wireheading problem. In this paper we suggest an alternative to RL called value reinforcement learning (VRL). In VRL, agents use the reward signal to learn a utility function. The VRL setup allows us to remove the incentive to wirehead by placing a constraint on the agent's actions. The constraint is defined in terms of the agent's belief distributions, and does not require an explicit specification of which actions constitute wireheading.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI Safety for Everyone

    cs.CY 2025-02 conditional novelty 5.0 of 10

    A systematic review of 383 papers argues that AI safety research already covers a wide spectrum of concrete, near-term concerns and should be understood as part of traditional technological safety practice.

Pith tools