Pith. sign in

REVIEW

Restricted Value Iteration: Theory and Algorithms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1107.0042 v1 pith:WJX4DD4Y submitted 2011-06-30 cs.AI

classification cs.AI
keywords iterationvaluebeliefpomdpsrestrictedpoliciesspacesubsets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Value iteration is a popular algorithm for finding near optimal policies for POMDPs. It is inefficient due to the need to account for the entire belief space, which necessitates the solution of large numbers of linear programs. In this paper, we study value iteration restricted to belief subsets. We show that, together with properly chosen belief subsets, restricted value iteration yields near-optimal policies and we give a condition for determining whether a given belief subset would bring about savings in space and time. We also apply restricted value iteration to two interesting classes of POMDPs, namely informative POMDPs and near-discernible POMDPs.

Discussion (0). Continue with ORCID to comment.

Pith tools