Pith. sign in

REVIEW 1 cited by

Automated Feature Selection for Inverse Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.15079 v1 pith:FH4NLTNA submitted 2024-03-22 cs.LG cs.RO

classification cs.LGcs.RO
keywords learningfeaturesrewardfeaturefunctionsreinforcementstateapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Inverse reinforcement learning (IRL) is an imitation learning approach to learning reward functions from expert demonstrations. Its use avoids the difficult and tedious procedure of manual reward specification while retaining the generalization power of reinforcement learning. In IRL, the reward is usually represented as a linear combination of features. In continuous state spaces, the state variables alone are not sufficiently rich to be used as features, but which features are good is not known in general. To address this issue, we propose a method that employs polynomial basis functions to form a candidate set of features, which are shown to allow the matching of statistical moments of state distributions. Feature selection is then performed for the candidates by leveraging the correlation between trajectory probabilities and feature expectations. We demonstrate the approach's effectiveness by recovering reward functions that capture expert policies across non-linear control tasks of increasing complexity. Code, data, and videos are available at https://sites.google.com/view/feature4irl.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Active Probing with Multimodal Predictions for Motion Planning

    cs.RO 2025-07 conditional novelty 6.0 of 10

    An MPC framework that uses a Wasserstein-based risk metric and a Boltzmann model of agent behavior to actively probe and infer other vehicles' intentions in multimodal prediction settings.

Pith tools