Pith. sign in

REVIEW 2 cited by

An Analysis of Frame-skipping in Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.03718 v1 pith:A54I45WM submitted 2021-02-07 cs.LG

classification cs.LG
keywords learningsensingstepsactionaction-repetitionalgorithmsanalysisatari
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In the practice of sequential decision making, agents are often designed to sense state at regular intervals of $d$ time steps, $d > 1$, ignoring state information in between sensing steps. While it is clear that this practice can reduce sensing and compute costs, recent results indicate a further benefit. On many Atari console games, reinforcement learning (RL) algorithms deliver substantially better policies when run with $d > 1$ -- in fact with $d$ even as high as $180$. In this paper, we investigate the role of the parameter $d$ in RL; $d$ is called the "frame-skip" parameter, since states in the Atari domain are images. For evaluating a fixed policy, we observe that under standard conditions, frame-skipping does not affect asymptotic consistency. Depending on other parameters, it can possibly even benefit learning. To use $d > 1$ in the control setting, one must first specify which $d$-step open-loop action sequences can be executed in between sensing steps. We focus on "action-repetition", the common restriction of this choice to $d$-length sequences of the same action. We define a task-dependent quantity called the "price of inertia", in terms of which we upper-bound the loss incurred by action-repetition. We show that this loss may be offset by the gain brought to learning by a smaller task horizon. Our analysis is supported by experiments on different tasks and learning algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Select before Act: Spatially Decoupled Action Repetition for Continuous Control

    cs.LG 2025-02 conditional novelty 6.0 of 10

    SDAR performs closed-loop act-or-repeat selection separately for each action dimension, improving sample efficiency and reducing action fluctuation in continuous control.

  2. Bounding Distributional Shifts in World Modeling through Novelty Detection

    cs.RO 2025-08 conditional novelty 4.0 of 10

    Attaching a VAE novelty detector to the DINO-WM world model and penalizing out-of-distribution predicted states in CEM planning lowers Chamfer distance on small-data robot manipulation benchmarks.

Pith tools