Pith. sign in

REVIEW 1 cited by

Performative Reinforcement Learning in Gradually Shifting Environments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09838 v2 pith:MW64X7SS submitted 2024-02-15 cs.LG

classification cs.LG
keywords environmentmdrralgorithmsdeployeddynamicsperformativepreviousalgorithm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When Reinforcement Learning (RL) agents are deployed in practice, they might impact their environment and change its dynamics. We propose a new framework to model this phenomenon, where the current environment depends on the deployed policy as well as its previous dynamics. This is a generalization of Performative RL (PRL) [Mandal et al., 2023]. Unlike PRL, our framework allows to model scenarios where the environment gradually adjusts to a deployed policy. We adapt two algorithms from the performative prediction literature to our setting and propose a novel algorithm called Mixed Delayed Repeated Retraining (MDRR). We provide conditions under which these algorithms converge and compare them using three metrics: number of retrainings, approximation guarantee, and number of samples per deployment. MDRR is the first algorithm in this setting which combines samples from multiple deployments in its training. This makes MDRR particularly suitable for scenarios where the environment's response strongly depends on its previous dynamics, which are common in practice. We experimentally compare the algorithms using a simulation-based testbed and our results show that MDRR converges significantly faster than previous approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Simulating Eutopia: Revisiting Long-term Fairness with Outcomes, Performativity, and Dynamics

    cs.CY 2026-07 conditional novelty 6.0 of 10

    In a performative lending simulator, modeling policy-induced distribution shift and optimizing outcome-based fairness rewards reduces long-run wealth inequality without sacrificing profit.

Pith tools