Pith. sign in

REVIEW 1 cited by

Reinforcement Learning in Presence of Discrete Markovian Context Evolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.06557 v1 pith:CZ4272KV submitted 2022-02-14 cs.LG cs.AI

classification cs.LGcs.AI
keywords contextlearningcontextsapproachargueevolutionmarkoviannumber
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider a context-dependent Reinforcement Learning (RL) setting, which is characterized by: a) an unknown finite number of not directly observable contexts; b) abrupt (discontinuous) context changes occurring during an episode; and c) Markovian context evolution. We argue that this challenging case is often met in applications and we tackle it using a Bayesian approach and variational inference. We adapt a sticky Hierarchical Dirichlet Process (HDP) prior for model learning, which is arguably best-suited for Markov process modeling. We then derive a context distillation procedure, which identifies and removes spurious contexts in an unsupervised fashion. We argue that the combination of these two components allows to infer the number of contexts from data thus dealing with the context cardinality assumption. We then find the representation of the optimal policy enabling efficient policy learning using off-the-shelf RL algorithms. Finally, we demonstrate empirically (using gym environments cart-pole swing-up, drone, intersection) that our approach succeeds where state-of-the-art methods of other frameworks fail and elaborate on the reasons for such failures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning What Matters Now: A Dual-Critic Context-Aware RL Framework for Priority-Driven Information Gain

    cs.AI 2025-06 conditional novelty 6.0 of 10

    CA-MIQ combines a novelty-driven intrinsic critic with a priority-shift detector and selective value resets, achieving about four times higher mission success than baseline Q-learning after priority changes in a simul...

Pith tools