Pith. sign in

REVIEW 1 cited by

Apprenticeship Learning for Model Parameters of Partially Observable Environments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1206.6484 v1 pith:GTVDPUUL submitted 2012-06-27 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords environmentmodelapprenticeshipdemonstrationexpertlearningobservableparameters
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider apprenticeship learning, i.e., having an agent learn a task by observing an expert demonstrating the task in a partially observable environment when the model of the environment is uncertain. This setting is useful in applications where the explicit modeling of the environment is difficult, such as a dialogue system. We show that we can extract information about the environment model by inferring action selection process behind the demonstration, under the assumption that the expert is choosing optimal actions based on knowledge of the true model of the target environment. Proposed algorithms can achieve more accurate estimates of POMDP parameters and better policies from a short demonstration, compared to methods that learns only from the reaction from the environment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Inverse Reinforcement Learning using Revealed Preferences and Passive Stochastic Optimization

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A three-chapter monograph that uses Afriat's theorem and Bayesian revealed preference tests for inverse reinforcement learning, plus a passive Langevin dynamics algorithm for real-time reward reconstruction.

Pith tools