Pith. sign in

REVIEW 1 cited by

Act-Then-Measure: Reinforcement Learning for Partially Observable Environments with Active Measuring

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.08271 v1 pith:SNLEYIHS submitted 2023-03-14 cs.AI cs.LG

classification cs.AIcs.LG
keywords heuristicactioncontrolobservableact-then-measureactionsaffectsenvironments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study Markov decision processes (MDPs), where agents have direct control over when and how they gather information, as formalized by action-contingent noiselessly observable MDPs (ACNO-MPDs). In these models, actions consist of two components: a control action that affects the environment, and a measurement action that affects what the agent can observe. To solve ACNO-MDPs, we introduce the act-then-measure (ATM) heuristic, which assumes that we can ignore future state uncertainty when choosing control actions. We show how following this heuristic may lead to shorter policy computation times and prove a bound on the performance loss incurred by the heuristic. To decide whether or not to take a measurement action, we introduce the concept of measuring value. We develop a reinforcement learning algorithm based on the ATM heuristic, using a Dyna-Q variant adapted for partially observable domains, and showcase its superior performance compared to prior methods on a number of partially-observable environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. To Measure or Not: A Cost-Sensitive, Selective Measuring Environment for Agricultural Management Decisions with Reinforcement Learning

    cs.LG 2025-01 conditional novelty 5.0 of 10

    A cost-sensitive reinforcement learning environment shows an agent can learn when to pay for crop measurements to guide nitrogen fertilization in winter wheat.

Pith tools