Pith. sign in

REVIEW 1 cited by

Counting Reward Automata: Sample Efficient Reinforcement Learning Through the Exploitation of Reward Function Structure

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.11364 v2 pith:UWEEBRSO submitted 2023-12-18 cs.AI

classification cs.AI
keywords rewardapproachesautomatonlanguagesampletaskscomplexitycounting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present counting reward automata-a finite state machine variant capable of modelling any reward function expressible as a formal language. Unlike previous approaches, which are limited to the expression of tasks as regular languages, our framework allows for tasks described by unrestricted grammars. We prove that an agent equipped with such an abstract machine is able to solve a larger set of tasks than those utilising current approaches. We show that this increase in expressive power does not come at the cost of increased automaton complexity. A selection of learning algorithms are presented which exploit automaton structure to improve sample efficiency. We show that the state machines required in our formulation can be specified from natural language task descriptions using large language models. Empirical results demonstrate that our method outperforms competing approaches in terms of sample efficiency, automaton complexity, and task completion.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning

    cs.LG 2026-08 conditional novelty 7.0 of 10

    An identical exploration bonus amplifies, equalizes, or has no effect on memory architectures depending on whether the task requires active discovery, a single reward-supervised cue, or follows a fixed schedule.

Pith tools