Pith. sign in

REVIEW 1 cited by

Estimation of Optimal Dynamic Treatment Assignment Rules under Policy Constraints

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.05031 v5 pith:5P5CUUG3 submitted 2021-06-09 econ.EM stat.MEstat.ML

classification econ.EMstat.MEstat.ML
keywords treatmentoptimaldynamicestimationassignmentconstraintsmethodspolicies
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Many policies involve dynamics in their treatment assignments, where individuals receive sequential interventions over multiple stages. We study estimation of an optimal dynamic treatment regime that guides the optimal treatment assignment for each individual at each stage based on their history. We propose an empirical welfare maximization approach in this dynamic framework, which estimates the optimal dynamic treatment regime using data from an experimental or quasi-experimental study while satisfying exogenous constraints on policies. The paper proposes two estimation methods: one solves the treatment assignment problem sequentially through backward induction, and the other solves the entire problem simultaneously across all stages. We establish finite-sample upper bounds on worst-case average welfare regrets for these methods and show their optimal $n^{-1/2}$ convergence rates. We also modify the simultaneous estimation method to accommodate intertemporal budget/capacity constraints.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Who With Whom? Learning Optimal Matching Policies

    econ.EM 2025-07 conditional novelty 6.0 of 10

    An entropy-regularized optimal transport method learns welfare-optimal two-sided matching policies with estimated costs, supported by a non-asymptotic regret bound and calibrated simulations suggesting about one perce...

Pith tools