Pith. sign in

REVIEW 1 cited by

Fairness in Reinforcement Learning: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.06909 v1 pith:GL7HS7LS submitted 2024-05-11 cs.LG cs.AIcs.CY

classification cs.LGcs.AIcs.CY
keywords fairnesssystemsbeenlearningunderstandingfairliteraturereal-world
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While our understanding of fairness in machine learning has significantly progressed, our understanding of fairness in reinforcement learning (RL) remains nascent. Most of the attention has been on fairness in one-shot classification tasks; however, real-world, RL-enabled systems (e.g., autonomous vehicles) are much more complicated in that agents operate in dynamic environments over a long period of time. To ensure the responsible development and deployment of these systems, we must better understand fairness in RL. In this paper, we survey the literature to provide the most up-to-date snapshot of the frontiers of fairness in RL. We start by reviewing where fairness considerations can arise in RL, then discuss the various definitions of fairness in RL that have been put forth thus far. We continue to highlight the methodologies researchers used to implement fairness in single- and multi-agent RL systems before showcasing the distinct application domains that fair RL has been investigated in. Finally, we critically examine gaps in the literature, such as understanding fairness in the context of RLHF, that still need to be addressed in future work to truly operationalize fair RL in real-world systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fairness Aware Reinforcement Learning via Proximal Policy Optimization

    cs.MA 2025-02 conditional novelty 5.0 of 10

    Adding retrospective and prospective reward-disparity penalties to PPO lowers demographic parity and conditional statistical parity disparities in two multi-agent simulations, at a measurable efficiency cost.

Pith tools