Pith. sign in

REVIEW

Active Reinforcement Learning: Observing Rewards at a Cost

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.06709 v2 pith:KPW6H4XH submitted 2020-11-13 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningreinforcementactivebanditscostinformationmulti-armedreward
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Active reinforcement learning (ARL) is a variant on reinforcement learning where the agent does not observe the reward unless it chooses to pay a query cost c > 0. The central question of ARL is how to quantify the long-term value of reward information. Even in multi-armed bandits, computing the value of this information is intractable and we have to rely on heuristics. We propose and evaluate several heuristic approaches for ARL in multi-armed bandits and (tabular) Markov decision processes, and discuss and illustrate some challenging aspects of the ARL problem.

Discussion (0). Continue with ORCID to comment.

Pith tools