Pith. sign in

REVIEW 1 cited by

Maximum Likelihood Constraint Inference for Inverse Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.05477 v2 pith:O3E4ABA6 submitted 2019-09-12 cs.LG cs.AIcs.ROcs.SYeess.SYstat.ML

classification cs.LGcs.AIcs.ROcs.SYeess.SYstat.ML
keywords behavioragentconstraintslikelihoodgivenmaximumrewardbest
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While most approaches to the problem of Inverse Reinforcement Learning (IRL) focus on estimating a reward function that best explains an expert agent's policy or demonstrated behavior on a control task, it is often the case that such behavior is more succinctly represented by a simple reward combined with a set of hard constraints. In this setting, the agent is attempting to maximize cumulative rewards subject to these given constraints on their behavior. We reformulate the problem of IRL on Markov Decision Processes (MDPs) such that, given a nominal model of the environment and a nominal reward function, we seek to estimate state, action, and feature constraints in the environment that motivate an agent's behavior. Our approach is based on the Maximum Entropy IRL framework, which allows us to reason about the likelihood of an expert agent's demonstrations given our knowledge of an MDP. Using our method, we can infer which constraints can be added to the MDP to most increase the likelihood of observing these demonstrations. We present an algorithm which iteratively infers the Maximum Likelihood Constraint to best explain observed behavior, and we evaluate its efficacy using both simulated behavior and recorded data of humans navigating around an obstacle.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DIAL: Distribution-Informed Adaptive Learning of Multi-Task Constraints for Safety-Critical Systems

    cs.LG 2025-01 conditional novelty 6.0 of 10

    DIAL learns a Beta-distributed safety constraint model from multi-task demonstrations and adapts it to new tasks via a tuned CVaR risk level, improving safety in RL transfer benchmarks.

Pith tools