Pith. sign in

REVIEW 2 cited by

Walking the Values in Bayesian Inverse Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.10971 v1 pith:S56HPH5K submitted 2024-07-15 cs.LG

classification cs.LG
keywords bayesianrewardsrewardspacevaluescarlocomputationinverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The goal of Bayesian inverse reinforcement learning (IRL) is recovering a posterior distribution over reward functions using a set of demonstrations from an expert optimizing for a reward unknown to the learner. The resulting posterior over rewards can then be used to synthesize an apprentice policy that performs well on the same or a similar task. A key challenge in Bayesian IRL is bridging the computational gap between the hypothesis space of possible rewards and the likelihood, often defined in terms of Q values: vanilla Bayesian IRL needs to solve the costly forward planning problem - going from rewards to the Q values - at every step of the algorithm, which may need to be done thousands of times. We propose to solve this by a simple change: instead of focusing on primarily sampling in the space of rewards, we can focus on primarily working in the space of Q-values, since the computation required to go from Q-values to reward is radically cheaper. Furthermore, this reversion of the computation makes it easy to compute the gradient allowing efficient sampling using Hamiltonian Monte Carlo. We propose ValueWalk - a new Markov chain Monte Carlo method based on this insight - and illustrate its advantages on several tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Distributional Inverse Reinforcement Learning

    cs.LG 2025-10 reject novelty 6.0 of 10

    DistIRL recovers reward distributions and risk-aware policies from offline demonstrations by minimizing first-order stochastic dominance violations between agent and expert returns.

  2. Distributional Inverse Reinforcement Learning

    cs.LG 2025-10 unverdicted novelty 5.0 of 10

    A distributional offline IRL method minimizes first-order stochastic dominance violations to recover reward distributions and distribution-aware policies, with O(ε^{-2}) convergence and reported SOTA results on synthe...

Pith tools