Pith. sign in

REVIEW 2 cited by

Policy Learning for Optimal Dynamic Treatment Regimes with Observational Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00221 v7 pith:S5IFFVTT submitted 2024-03-30 stat.ME econ.EMmath.STstat.MLstat.TH

classification stat.MEecon.EMmath.STstat.MLstat.TH
keywords optimaltreatmentdynamiclearningpolicyconvergencedatainterventions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Public policies and medical interventions often involve dynamic treatment assignments, in which individuals receive a sequence of interventions over multiple stages. We study the statistical learning of optimal dynamic treatment regimes (DTRs) that determine the optimal treatment assignment for each individual at each stage based on their evolving history. We propose a novel, doubly robust, classification-based method for learning the optimal DTR from observational data under the sequential ignorability assumption. The method proceeds via backward induction: at each stage, it constructs and maximizes an augmented inverse probability weighting (AIPW) estimator of the policy value function to learn the optimal stage-specific policy. We show that the resulting DTR achieves an optimal convergence rate of $n^{-1/2}$ for welfare regret under mild convergence conditions on estimators of the nuisance components.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Minimax and Bayes Optimal Best-Arm Identification

    econ.EM 2025-06 conditional novelty 8.0 of 10

    TS-SPAS attains the exact asymptotic minimax and Bayes constants for fixed-budget best-arm identification, with matching lower and upper bounds over exponential family outcomes.

  2. Evaluating Program Sequences with Double Machine Learning: An Application to Labor Market Policies

    econ.EM 2025-06 conditional novelty 6.0 of 10

    A two-period double machine learning evaluation of Swiss labor market programs finds temporary wage subsidies the most effective programs when modeled as dynamic policies.

Pith tools