Pith. sign in

REVIEW 2 cited by

Biased Aggregation, Rollout, and Enhanced Policy Improvement for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.02426 v1 pith:GVJ45LT6 submitted 2019-10-06 cs.LG cs.AI

Biased Aggregation, Rollout, and Enhanced Policy Improvement for Reinforcement Learning

classification cs.LG cs.AI
keywords whenaggregatefunctionpolicyaggregationapproximateschemeapproximation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We propose a new aggregation framework for approximate dynamic programming, which provides a connection with rollout algorithms, approximate policy iteration, and other single and multistep lookahead methods. The central novel characteristic is the use of a bias function $V$ of the state, which biases the values of the aggregate cost function towards their correct levels. The classical aggregation framework is obtained when $V\equiv0$, but our scheme works best when $V$ is a known reasonably good approximation to the optimal cost function $J^*$. When $V$ is equal to the cost function $J_{\mu}$ of some known policy $\mu$ and there is only one aggregate state, our scheme is equivalent to the rollout algorithm based on $\mu$ (i.e., the result of a single policy improvement starting with the policy $\mu$). When $V=J_{\mu}$ and there are multiple aggregate states, our aggregation approach can be used as a more powerful form of improvement of $\mu$. Thus, when combined with an approximate policy evaluation scheme, our approach can form the basis for a new and enhanced form of approximate policy iteration. When $V$ is a generic bias function, our scheme is equivalent to approximation in value space with lookahead function equal to $V$ plus a local correction within each aggregate state. The local correction levels are obtained by solving a low-dimensional aggregate DP problem, yielding an arbitrarily close approximation to $J^*$, when the number of aggregate states is sufficiently large. Except for the bias function, the aggregate DP problem is similar to the one of the classical aggregation framework, and its algorithmic solution by simulation or other methods is nearly identical to one for classical aggregation, assuming values of $V$ are available when needed.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. An Error Bound for Aggregation in Approximate Dynamic Programming

    math.OC 2025-07 unverdicted novelty 5.0

    Generalizes the Tsitsiklis-van Roy error bound for aggregation in discounted DP to soft and feature-based schemes.

  2. Adaptive Network Security Policies via Belief Aggregation and Rollout

    eess.SY 2025-07 unverdicted novelty 4.0

    Combines particle filtering, feature-based aggregation, and rollout to produce scalable network security policies with theoretical guarantees that adapt quickly to model changes.