Pith. sign in

Policy Learning with Competing Agents

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Decision makers often aim to learn a treatment assignment policy under a capacity constraint on the number of agents that they can treat. When agents can respond strategically to such policies, competition arises, complicating estimation of the optimal policy. In this paper, we study capacity-constrained treatment assignment in the presence of such interference. We consider a dynamic model where the decision maker allocates treatments at each time step and heterogeneous agents myopically best respond to the previous treatment assignment policy. When the number of agents is large but finite, we show that the threshold for receiving treatment under a given policy converges to the policy's mean-field equilibrium threshold. Based on this result, we develop a consistent estimator for the policy gradient. In a semi-synthetic experiment with data from the National Education Longitudinal Study of 1988, we demonstrate that this estimator can be used for learning capacity-constrained policies in the presence of strategic behavior.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2024 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

support 1

representative citing papers

Practical Performative Policy Learning with Strategic Agents

cs.LG · 2024-12-02 · conditional · novelty 6.0

A strategic policy gradient method uses a low-dimensional evaluation vector of the deployed policy to learn agents' responses with a differentiable classifier, replacing parametric distribution maps.

citing papers explorer

Showing 1 of 1 citing paper.

  • Practical Performative Policy Learning with Strategic Agents cs.LG · 2024-12-02 · conditional · none · ref 57 · internal anchor

    A strategic policy gradient method uses a low-dimensional evaluation vector of the deployed policy to learn agents' responses with a differentiable classifier, replacing parametric distribution maps.