Pith. sign in

REVIEW 1 cited by

Synthesis of Discounted-Reward Optimal Policies for Markov Decision Processes Under Linear Temporal Logic Specifications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.00632 v2 pith:LUYEGKDC submitted 2020-11-01 eess.SY cs.FLcs.SY

classification eess.SYcs.FLcs.SY
keywords rewardunderlinearobjectiveoptimalpolicydecisiondiscounted
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a method to find an optimal policy with respect to a reward function for a discounted Markov decision process under general linear temporal logic (LTL) specifications. Previous work has either focused on maximizing a cumulative reward objective under finite-duration tasks, specified by syntactically co-safe LTL, or maximizing an average reward for persistent (e.g., surveillance) tasks. This paper extends and generalizes these results by introducing a pair of occupancy measures to express the LTL satisfaction objective and the expected discounted reward objective, respectively. These occupancy measures are then connected to a single policy via a novel reduction resulting in a mixed integer linear program whose solution provides an optimal policy. Our formulation can also be extended to include additional constraints with respect to secondary reward functions. We illustrate the effectiveness of our approach in the context of robotic motion planning for complex missions under uncertainty and performance objectives.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-agent Path Finding for Timed Tasks using Evolutionary Games

    cs.MA 2024-11 conditional novelty 4.0 of 10

    MAPF-EGT, an evolutionary game theory policy learner with weighted automaton rewards, is reported to beat deep RL baselines on large-grid multi-agent pathfinding.

Pith tools