REVIEW 1 major objections 5 minor 102 references
Optimal Targeting in Dynamic Systems
T0 review · 1 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Under mild assumptions, the optimal treatment policy in a dynamic system is a state-specific threshold on the direct treatment effect, learnable by augmenting CATE estimation with state-level value iteration.
desk verdict Theorem 2's threshold form is a clean Bellman-equation consequence and the SACT algorithm is a sensible practical bridge between CATE targeting and dynamic capacity, but Section 3's reward-rate extension has a real indexing flaw in the current draft. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is the state-level Bellman operator obtained by integrating out the exogenous covariates. Because $X_i$ is drawn from a fixed distribution $P_X$ independent of state and treatment, the continuation value collapses to a relative value function $V^*$ on the finite state space $\mathcal{S}$, and the covariate enters only through the CADE $\tau(x,s)$. This yields the operator $(T_{\mathcal{S}}v)(s)=r_0(s)+(P_{S,0}v)(s)+\mathbb{E}_X[(\tau(X,s)+((P_{S,1}-P_{S,0})v)(s))_+]$, where $r_0(s)=\mathbb{E}_X[\eta_0(X,s)]$ is the average baseline reward and $P_{S,w}(s'|s)$ is the state transition kernel under treatment $w$. Algorithm 1 estimates $\tau$, $r_0$, and the transition kernels on one sample split and applies relative value iteration to an empirical version of $T_{\mathcal{S}}$; the fixed point $V^*$ supplies the thresholds $c_s$. The same operator is adapted to reward-rate objectives by replacing the outcome with $R_i-\theta^*\Delta_i$ and iterating via Dinkelbach's method.
What would settle it
Simulate the emergency-department parallel-queue system but generate patient covariates from a distribution that depends on the current fast-track queue length (or on past routing decisions), violating Assumption 1. If an exhaustive-search optimal policy strictly beats the CADE-threshold rule with the true state-specific thresholds computed from the true $V^*$, the characterization fails exactly under its stated boundary; equivalently, checking whether the estimated covariate distribution $P_X$ varies with $S$ in real queuing logs would test the premise.
Extended reading notes
Core claim
The central claim is Theorem 2: under time-homogeneous dynamics with exogenous covariates (Assumption 1) and a finite irreducible, aperiodic state process (Assumption 2), there exists an optimal deterministic policy of the form $\pi^*(x,s)=\mathbb{I}(\tau(x,s)>c_s)$, where $c_s = -\sum_{s'\in\mathcal{S}}(P_{S,1}(s'|s)-P_{S,0}(s'|s))V^*(s')$ and $V^*$ is the relative value function of the state-level average-reward Bellman equation. Thus the optimal dynamic targeting rule keeps the familiar CADE ranking within each state but changes the treatment cutoff, and the threshold is the dynamic shadow cost of assigning treatment in that state at the optimum. The direct targeting rule $\mathbb{I}(\tau(x,s)>0)$ is optimal only in the degenerate case where treatment does not affect state transitions. The paper supports this with regret bounds showing that the learned threshold policy has regret $O_P(n^{-(\beta\wedge 1/2)})$ when the CADE estimator achieves $L^2$ rate $n^{-\beta}$.
Load-bearing premise
The load-bearing premise is Assumption 1: each arriving unit's covariates are drawn from the same fixed distribution, independent of the system state and of past treatments; if the mix of arriving units shifts with congestion or with the policy, the state-level Bellman equation and the threshold characterization do not hold.
Editorial extensions
If this is right
- Existing CATE/CADE estimators (e.g., causal forests) can be reused as-is; the dynamic adjustment is confined to a low-dimensional state-level value iteration.
- Under congestion, direct targeting is suboptimal: it admits marginal patients who crowd out later high-benefit arrivals, and the gap grows with the strength of indirect effects.
- The threshold $c_s$ equals the negative conditional average indirect effect at the optimal policy, so the optimal rule is a direct/indirect effect decomposition: treat iff CADE exceeds the shadow cost.
- For systems where arrivals respond to congestion, the same threshold characterization holds for the long-run reward rate, using transformed outcome $R_i - \theta^* \Delta_i$ and Dinkelbach updates.
- In tabular settings, exploiting the exogenous-covariate structure yields regret smaller than the collapsed-state Markov-decision-process baseline by a factor of order $\min\{(t_0+t_{\mathrm{mix}})\sqrt{|\mathcal{X}||\mathcal{S}|}, |\mathcal{X}|\}$, making the method attractive with moderately large covariate spaces.
Reading between the lines
- One can read $c_s$ as a state-indexed shadow price; this suggests that related resource-allocation problems with convex congestion costs might admit analogous threshold rules, a direction the paper does not pursue.
- A practical diagnostic suggested by the theory: estimate CADE and CAIE separately; a near-zero CAIE means direct targeting is near optimal, while a large CAIE indicates capacity effects dominate.
- The regret analysis leaves margin conditions untouched; adding the usual margin assumption near the boundaries $\tau(x,s)=c_s$ could plausibly sharpen the $n^{-(\beta\wedge 1/2)}$ bound to faster plug-in rates, though the paper does not claim this.
- The finite-state assumption is used for the mixing and value bounds; extending SACT to continuous states would require function approximation for $V^*$ and a new analysis of the state-level operator.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies optimal treatment targeting in a Markovian system in which units arrive sequentially, each with covariates X_i, and the decision-maker assigns treatment W_i based on X_i and a system state S_i. The outcome Y_i may depend on the current state and action, and the action also affects the future state. Under a time-homogeneous MDP with exogenous covariates (Assumption 1), the paper proves (Theorem 2) that an optimal average-reward policy exists and takes the form π*(x,s)=I(τ(x,s)>c_s), where τ is the CADE and c_s = −Σ_{s'}(P_{S,1}(s'|s)−P_{S,0}(s'|s))V*(s') is a state-dependent shadow cost equal to the negative conditional average indirect effect. It proposes Algorithm 1 (SACT), which estimates τ via causal ML and V* via state-level relative value iteration, with sample-splitting by regenerative blocks. Regret bounds (Theorem 4) and a tabular comparison with collapsed-state RL (Proposition 5) are given, plus simulations in emergency-department routing and online customer support with congestion-sensitive arrivals, using a Dinkelbach reward-rate extension in Section 3.
Significance. The Section 2 characterization is an elegant and potentially practically important bridge between CATE-based targeting and dynamic programming: it reduces a high-dimensional MDP with covariates to a low-dimensional state-level Bellman recursion, with a decision rule that preserves the familiar CADE ranking within each state. The regret analysis and the finite-state comparison provide a credible argument for a statistical advantage over collapsed-state RL, and the numerical studies show large gains over direct targeting and generic offline RL. The main weakness is the reward-rate extension in Section 3, which is conceptually incorrect as written; because this extension underlies the customer-support experiment, the paper's claims about congestion-sensitive arrivals require revision. The core Section 2 theorem, however, appears sound.
major comments (1)
- [§3, Eqs. (26)-(29), Section 4.2] The reward-rate extension is built on the transformed outcome Ỹ_i = R_i − θ*Δ_i, where Δ_i = T_i − T_{i−1} is the inter-event time preceding event i. Because Δ_i is realized before the action W_i is chosen, its conditional distribution given the current state and action is independent of W_i, so τ_Δ(x,s) = 0 identically. Consequently, the Bellman operator (29) does not charge the decision for the effect of treatment on future arrival times: the sojourn time associated with the decision is Δ_{i+1}, not Δ_i. The correct Dinkelbach inner problem for the semi-Markov reward rate θ(π) = E_π[R_i]/E_π[Δ_i] uses the forward inter-event time (e.g., Ỹ_i = R_i − θ*Δ_{i+1} or an equivalent reindexing), whose CADE is generically nonzero and belongs in the threshold. As written, the threshold rule (30) and the update (31) do not follow from the Section 2 theorem; the r_Δ,0(s) term in (31) is estimated under the behavior policy and is not the conditional mean backward recurrence time under the candidate policy π^(b), so the Dinkelbach iteration is also internally inconsistent. This undermines the theoretical claim of Section 3 and the reward-rate interpretation of the customer-support experiment in Section 4.2.
minor comments (5)
- [§2.2 and Appendix C.2.1] The regenerative-block sample-splitting procedure is described only informally; a precise construction of the blocks (choice of regeneration state, block lengths, how cross-fitting assigns whole blocks to folds) would make the algorithm and the statements in Footnote 4 reproducible.
- [Assumption 3] The uniform mixing assumption is imposed on the empirical kernel P̂_{S,π} as well as the true kernel, which is a strong condition not implied by Assumptions 1-2; the paper should clarify whether Assumption 3 is a high-level assumption the user must verify, and how the anchored-kernel suggestion of Zurek and Chen [2024] is incorporated into Algorithm 1.
- [Figure 4 caption] The caption refers to an 'effective state-specific threshold' but the figure displays a single horizontal line for the direct rule and the optimal rule is shown through the treated/untreated coloring; please make the display of the state-dependent thresholds under the optimal rule explicit.
- [Section A.1 and B.1] The symbol k is used both for the capacity bound and as a generic queue-length state, which is confusing (e.g., in Eq. (S3) the condition k < k appears); please introduce separate notation for the capacity limit and the running state.
- [Eq. (10)] The expression for the Bellman operator (10) uses E_X[(τ(X,s)+((P_{S,1}−P_{S,0})v)(s))^+], which is correct, but the subsequent presentation in Eq. (13) uses a sample average over the evaluation set without restating that this approximates the covariate expectation; making this explicit in the text would help readers.
Circularity Check
Reward-rate extension in Section 3 is self-definitional: Δ_i is realized before W_i, making τ_Δ≡0, so the claimed congestion-aware threshold reduces by construction to the Section 2 direct-reward operator.
-
self definitional
[Section 3, Eqs. (26)-(30), especially Eq. (29)]
"Let Δ_i :=T_i −T_{i−1} denote the inter-event time. ... This event-level representation has the same Markov structure as the model in Section 2. In particular, for any fixed scalar outcome constructed from (R_i,Δ_i), such as the transformed outcome introduced below, the state-aware Bellman characterization can be applied with state S_i = (K_i,A_i). ... τ_Ỹ(x,s)=τ_R(x,s)−θ*τ_Δ(x,s)."
Δ_i is the time before event i and is realized before W_i is chosen; conditionally on (X_i,S_i,W_i), its law does not depend on W_i, so the CADE τ_Δ(x,s) is identically zero. Therefore τ_Ỹ = τ_R, and the Bellman operator (29) coincides with the direct-reward operator of Section 2. The claimed state-specific thresholds for congestion-sensitive arrivals thus cannot encode the future arrival-rate channel the section introduces; the reward-rate extension reduces by construction to the earlier direct-reward analysis rather than deriving the arrival-rate-aware policy it announces. Additionally, Δ_i's law depends on S_{i−1} and the holding time before T_i, so Assumption 1's reward structure P_Y(·|X_i,S_i,W_i) is violated for Ỹ_i.
full rationale
The main result of Section 2, Theorem 2, is a genuine fixed-point characterization: the threshold c_s is defined from V*, which solves the average-reward Bellman equation, and the proof in Appendix D.1 does not assume the threshold form. The CADE/CAIE decomposition is derived from the policy-gradient theorem rather than imposed, and the self-citations to Hu et al. (2022), Munro et al. (2025), and Li and Wager (2022) are contextual and not load-bearing. However, the Section 3 extension is circular by construction. By defining Δ_i as the backward inter-event time T_i − T_{i−1}, the paper makes τ_Δ(x,s) identically zero, so the transformed outcome's CADE reduces to the direct reward CADE and the Bellman operator (29) is exactly the Section 2 operator with no congestion-arrival adjustment. The claimed reward-rate threshold (30) therefore does not follow from the stated model; it is equivalent, by construction, to the direct-reward analysis. This is a partial circularity affecting the reward-rate extension and the congestion-sensitive customer-support experiment, while the Section 2 threshold theorem remains independent and sound.
Assumptions & free parameters
assumptions (6)
- domain assumption Time-homogeneous MDP with exogenous covariates (Assumption 1).
- domain assumption Irreducibility, aperiodicity, finite state space, bounded outcomes (Assumption 2).
- ad hoc to paper Uniform mixing of both true and empirical transition kernels (Assumption 3).
- domain assumption Overlap and nuisance regularity for the doubly robust baseline estimator (Assumption 4).
- domain assumption Sequential ignorability: the status-quo policy depends only on (X_i, S_i).
- ad hoc to paper Regenerative blocks yield i.i.d. chunks for sample splitting.
Cite this review
Pith. "Pith review of Optimal Targeting in Dynamic Systems." pith.science (2026). https://pith.science/paper/B3KCNIEE
@misc{pith2026250700312,
author = {Pith},
title = {Pith review of: Optimal Targeting in Dynamic Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/B3KCNIEE}},
note = {Machine review of arXiv:2507.00312}
}
read the original abstract
Modern treatment targeting methods often rely on estimating a conditional average treatment effect (CATE) using machine learning tools. While effective in identifying who benefits from treatment on the individual level, these approaches typically overlook system-level dynamics that may arise when treatments induce strain on shared capacity. We study the problem of targeting in Markovian systems, where treatment decisions must be made one at a time as units arrive, and early decisions can impact later outcomes through delayed or limited access to resources. We show that optimal policies in such settings compare CATE-like quantities to state-specific thresholds, where each threshold reflects the expected cumulative impact on the system of treating an additional individual in the given state. We propose an algorithm that augments standard CATE estimation with state-level value iteration to estimate these thresholds from observational data. Theoretical results establish consistency and convergence guarantees, and empirical studies demonstrate that our method improves long-run outcomes considerably relative to individual-level CATE targeting rules and generic offline reinforcement learning algorithms.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Learning in structured mdps with convex cost functions: Improved regret bounds for inventory management
Shipra Agrawal and Randy Jia. Learning in structured mdps with convex cost functions: Improved regret bounds for inventory management. In Proceedings of the 2019 ACM Conference on Economics and Computation, pages 743--744, 2019
2019
-
[2]
Estimating average causal effects under general interference, with application to a social network experiment
Peter M Aronow and Cyrus Samii. Estimating average causal effects under general interference, with application to a social network experiment. The Annals of Applied Statistics, 11 0 (4): 0 1912--1947, 2017
1912
-
[3]
Recursive partitioning for heterogeneous causal effects
Susan Athey and Guido Imbens. Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences, 113 0 (27): 0 7353--7360, 2016
2016
-
[4]
Policy learning with observational data
Susan Athey and Stefan Wager. Policy learning with observational data. Econometrica, 89 0 (1): 0 133--161, 2021
2021
-
[5]
Generalized random forests
Susan Athey, Julie Tibshirani, and Stefan Wager. Generalized random forests. The Annals of Statistics, 47 0 (2): 0 1148--1178, 2019
2019
-
[6]
The big data newsvendor: Practical insights from machine learning
Gah-Yi Ban and Cynthia Rudin. The big data newsvendor: Practical insights from machine learning. Operations Research, 67 0 (1): 0 90--108, 2019
2019
-
[7]
Randomization tests of causal effects under interference
Guillaume W Basse, Avi Feller, and Panos Toulis. Randomization tests of causal effects under interference. Biometrika, 106 0 (2): 0 487--494, 2019
2019
-
[8]
From predictive to prescriptive analytics
Dimitris Bertsimas and Nathan Kallus. From predictive to prescriptive analytics. Management Science, 66 0 (3): 0 1025--1044, 2020
2020
Show all 102 references
-
[9]
Inferring welfare maximizing treatment assignment under budget constraints
Debopam Bhattacharya and Pascaline Dupas. Inferring welfare maximizing treatment assignment under budget constraints. Journal of Econometrics, 167 0 (1): 0 168--196, 2012
2012
-
[10]
The optimal admission threshold in observable queues with state dependent pricing
Christian Borgs, Jennifer T Chayes, Sherwin Doroudi, Mor Harchol-Balter, and Kuang Xu. The optimal admission threshold in observable queues with state dependent pricing. Probability in the Engineering and Informational Sciences, 28 0 (1): 0 101--119, 2014
2014
-
[11]
Randomized controlled trials of service interventions: The impact of capacity constraints
Justin Boutilier, Jonas Oddur Jonasson, Hannah Li, and Erez Yoeli. Randomized controlled trials of service interventions: The impact of capacity constraints. arXiv preprint arXiv:2407.21322, 2024
2024
-
[12]
Learning to order for inventory systems with lost sales and uncertain supplies
Boxiao Chen, Jiashuo Jiang, Jiawei Zhang, and Zhengyuan Zhou. Learning to order for inventory systems with lost sales and uncertain supplies. Management Science, 70 0 (12): 0 8631--8646, 2024
2024
-
[13]
State dependent pricing with a queue
Hong Chen and Murray Z Frank. State dependent pricing with a queue. IIE Transactions, 33 0 (10): 0 847--860, 2001
2001
-
[14]
A statistical learning approach to personalization in revenue management
Xi Chen, Zachary Owen, Clark Pixton, and David Simchi-Levi. A statistical learning approach to personalization in revenue management. Management Science, 68 0 (3): 0 1923--1937, 2022
1923
-
[15]
Double/debiased machine learning for treatment and structural parameters, 2018
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. Double/debiased machine learning for treatment and structural parameters, 2018
2018
-
[16]
Locally robust semiparametric estimation
Victor Chernozhukov, Juan Carlos Escanciano, Hidehiko Ichimura, Whitney K Newey, and James M Robins. Locally robust semiparametric estimation. Econometrica, 90 0 (4): 0 1501--1535, 2022
2022
-
[17]
Sampling-based approximation schemes for capacitated stochastic inventory control models
Wang Chi Cheung and David Simchi-Levi. Sampling-based approximation schemes for capacitated stochastic inventory control models. Mathematics of Operations Research, 44 0 (2): 0 668--692, 2019
2019
-
[18]
Comparison of perturbation bounds for the stationary distribution of a markov chain
Grace E Cho and Carl D Meyer. Comparison of perturbation bounds for the stationary distribution of a markov chain. Linear Algebra and its Applications, 335 0 (1-3): 0 137--150, 2001
2001
-
[19]
Discovering and removing exogenous state variables and rewards for reinforcement learning
Thomas Dietterich, George Trimponias, and Zhitang Chen. Discovering and removing exogenous state variables and rewards for reinforcement learning. In International Conference on Machine Learning, pages 1262--1270. PMLR, 2018
2018
-
[20]
Feature-based inventory control with censored demand
Jingying Ding, Woonghee Tim Huh, and Ying Rong. Feature-based inventory control with censored demand. Manufacturing & Service Operations Management, 26 0 (3): 0 1157--1172, 2024
2024
-
[21]
Sample-efficient reinforcement learning in the presence of exogenous information
Yonathan Efroni, Dylan J Foster, Dipendra Misra, Akshay Krishnamurthy, and John Langford. Sample-efficient reinforcement learning in the presence of exogenous information. In Conference on Learning Theory, pages 5062--5127. PMLR, 2022
2022
-
[22]
Markovian interference in experiments
Vivek Farias, Andrew Li, Tianyi Peng, and Andrew Zheng. Markovian interference in experiments. Advances in Neural Information Processing Systems, 35: 0 535--549, 2022
2022
-
[23]
Deep neural networks for estimation and inference
Max H Farrell, Tengyuan Liang, and Sanjog Misra. Deep neural networks for estimation and inference. Econometrica, 89 0 (1): 0 181--213, 2021
2021
-
[24]
How research in production and operations management may evolve in the era of big data
Qi Feng and J George Shanthikumar. How research in production and operations management may evolve in the era of big data. Production and Operations Management, 27 0 (9): 0 1670--1684, 2018
2018
-
[25]
Targeting interventions in networks
Andrea Galeotti, Benjamin Golub, and Sanjeev Goyal. Targeting interventions in networks. Econometrica, 88 0 (6): 0 2445--2471, 2020
2020
-
[26]
Causal Inference: What If
Miguel A Hern\'an and James M Robins. Causal Inference: What If. Chapman & Hall/CRC, Boca Raton, 2020
2020
-
[27]
Asymptotics for statistical treatment rules
Keisuke Hirano and Jack R Porter. Asymptotics for statistical treatment rules. Econometrica, 77 0 (5): 0 1683--1701, 2009
2009
-
[28]
Dynamic Programming and Markov Processes
Ronald A Howard. Dynamic Programming and Markov Processes. John Wiley and Sons/MIT Press, 1960
1960
-
[29]
Average direct and indirect causal effects under interference
Yuchen Hu, Shuangning Li, and Stefan Wager. Average direct and indirect causal effects under interference. Biometrika, 109 0 (4): 0 1165--1172, 2022
2022
-
[30]
Estimating individualized treatment rules with risk constraint
Xinyang Huang and Jin Xu. Estimating individualized treatment rules with risk constraint. Biometrics, 76 0 (4): 0 1310--1318, 2020
2020
-
[31]
Toward causal inference with interference
Michael G Hudgens and M Elizabeth Halloran. Toward causal inference with interference. Journal of the American Statistical Association, 103 0 (482): 0 832--842, 2008
2008
-
[32]
Resources and capabilities of the department of veterans affairs to provide timely and accessible care to veterans
Peter S Hussey, Jeanne S Ringel, Sangeeta Ahluwalia, Rebecca Anhang Price, Christine Buttorff, Thomas W Concannon, Susan L Lovejoy, Grant R Martsolf, Robert S Rudin, Dana Schultz, et al. Resources and capabilities of the department of veterans affairs to provide timely and acc...
2016
-
[33]
Experimental evaluation of individualized treatment rules
Kosuke Imai and Michael Lingzhi Li. Experimental evaluation of individualized treatment rules. Journal of the American Statistical Association, 118 0 (541): 0 242--256, 2023
2023
-
[34]
Priority assignment in emergency response
Evin Uzun Jacobson, Nilay Tan k Argon, and Serhan Ziya. Priority assignment in emergency response. Operations Research, 60 0 (4): 0 813--832, 2012
2012
-
[35]
Control of arrivals to a stochastic input--output system
S ren Glud Johansen and Shaler Stidham Jr. Control of arrivals to a stochastic input--output system. Advances in Applied Probability, 12 0 (4): 0 972--999, 1980
1980
-
[36]
When does interference matter? decision-making in platform experiments
Ramesh Johari, Hannah Li, Anushka Murthy, and Gabriel Y Weintraub. When does interference matter? decision-making in platform experiments. arXiv preprint arXiv:2410.06580, 2024
2024 arXiv
-
[37]
Steven G. Johnson. The NLopt nonlinear-optimization package, 2008. URL https://github.com/stevengj/nlopt
2008
-
[38]
Balanced policy evaluation and learning
Nathan Kallus. Balanced policy evaluation and learning. Advances in neural information processing systems, 31, 2018
2018
-
[39]
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Nathan Kallus and Masatoshi Uehara. Double reinforcement learning for efficient off-policy evaluation in markov decision processes. J. Mach. Learn. Res., 21: 0 167--1, 2020
2020
-
[40]
Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning
Nathan Kallus and Masatoshi Uehara. Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning. Operations Research, 70 0 (6): 0 3282--3302, 2022
2022
-
[41]
Minimax-optimal policy learning under unobserved confounding
Nathan Kallus and Angela Zhou. Minimax-optimal policy learning under unobserved confounding. Management Science, 67 0 (5): 0 2870--2890, 2021
2021
-
[42]
Towards optimal doubly robust estimation of heterogeneous causal effects
Edward H Kennedy. Towards optimal doubly robust estimation of heterogeneous causal effects. Electronic Journal of Statistics, 17 0 (2): 0 3008--3049, 2023
2023
-
[43]
Feature-based scheduling and dynamic learning with a large backlog
N Bora Keskin and Can Zhang. Feature-based scheduling and dynamic learning with a large backlog. Available at SSRN 4852356, 2024
2024
-
[44]
Who should be treated? empirical welfare maximization methods for treatment choice
Toru Kitagawa and Aleksey Tetenov. Who should be treated? empirical welfare maximization methods for treatment choice. Econometrica, 86 0 (2): 0 591--616, 2018
2018
-
[45]
Individualized treatment allocation in sequential network games
Toru Kitagawa and Guanyi Wang. Individualized treatment allocation in sequential network games. arXiv preprint arXiv:2302.05747, 2023 a
2023 arXiv
-
[46]
Who should get vaccinated? individualized allocation of vaccines over sir network
Toru Kitagawa and Guanyi Wang. Who should get vaccinated? individualized allocation of vaccines over sir network. Journal of Econometrics, 232 0 (1): 0 109--131, 2023 b
2023
-
[47]
o ren R K \
S \"o ren R K \"u nzel, Jasjeet S Sekhon, Peter J Bickel, and Bin Yu. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the national academy of sciences, 116 0 (10): 0 4156--4165, 2019
2019
-
[48]
Batch policy learning under constraints
Hoang Le, Cameron Voloshin, and Yisong Yue. Batch policy learning under constraints. In International Conference on Machine Learning, pages 3703--3712. PMLR, 2019
2019
-
[49]
Rate-optimal cluster-randomized designs for spatial interference
Michael P Leung. Rate-optimal cluster-randomized designs for spatial interference. The Annals of Statistics, 50 0 (5): 0 3064--3087, 2022
2022
-
[50]
Random graph asymptotics for treatment effect estimation under network interference
Shuangning Li and Stefan Wager. Random graph asymptotics for treatment effect estimation under network interference. arXiv preprint arXiv:2007.13302, 2020
2007 arXiv
-
[51]
Experimenting under stochastic congestion
Shuangning Li, Ramesh Johari, Xu Kuang, and Stefan Wager. Experimenting under stochastic congestion. arXiv preprint arXiv:2302.12093, 2023
2023
-
[52]
Off-policy estimation of long-term average outcomes with applications to mobile health
Peng Liao, Predrag Klasnja, and Susan Murphy. Off-policy estimation of long-term average outcomes with applications to mobile health. Journal of the American Statistical Association, 116 0 (533): 0 382--391, 2021
2021
-
[53]
Batch policy learning in average reward markov decision processes
Peng Liao, Zhengling Qi, Runzhe Wan, Predrag Klasnja, and Susan A Murphy. Batch policy learning in average reward markov decision processes. The Annals of Statistics, 50 0 (6): 0 3364--3387, 2022
2022
-
[54]
Performance guarantees for policy learning
Alex Luedtke and Antoine Chambaz. Performance guarantees for policy learning. Annales de l'IHP Probabilites et statistiques, 56 0 (3): 0 2162, 2020
2020
-
[55]
Optimal individualized treatments in resource-limited settings
Alexander R Luedtke and Mark J van der Laan. Optimal individualized treatments in resource-limited settings. The International Journal of Biostatistics, 12 0 (1): 0 283--303, 2016
2016
-
[56]
Charles F. Manski. Statistical treatment rules for heterogeneous populations. Econometrica, 72 0 (4): 0 1221--1246, 2004. doi:https://doi.org/10.1111/j.1468-0262.2004.00530.x
2004
-
[57]
Identification for prediction and decision
Charles F Manski. Identification for prediction and decision. Harvard University Press, 2009
2009
-
[58]
Identification of treatment response with social interactions
Charles F Manski. Identification of treatment response with social interactions. The Econometrics Journal, 16 0 (1): 0 S1--S23, 2013
2013
-
[59]
A vector-contraction inequality for rademacher complexities
Andreas Maurer. A vector-contraction inequality for rademacher complexities. In Algorithmic Learning Theory, pages 3--17, Cham, 2016. Springer International Publishing
2016
-
[60]
Off-policy evaluation in markov decision processes under weak distributional overlap
Mohammad Mehrabi and Stefan Wager. Off-policy evaluation in markov decision processes under weak distributional overlap. arXiv preprint arXiv:2402.08201, 2024
2024
-
[61]
Fast track: Urgent care within a teaching hospital emergency department: Can it work? Annals of Emergency Medicine, 17 0 (5): 0 453--456, 1988
Harvey W Meislin, Sally A Coates, Janine Cyr, and Terry Valenzuela. Fast track: Urgent care within a teaching hospital emergency department: Can it work? Annals of Emergency Medicine, 17 0 (5): 0 453--456, 1988
1988
-
[62]
Resource-based patient prioritization in mass-casualty incidents
Alex F Mills, Nilay Tan k Argon, and Serhan Ziya. Resource-based patient prioritization in mass-casualty incidents. Manufacturing & Service Operations Management, 15 0 (3): 0 361--377, 2013
2013
-
[63]
Treatment effects in market equilibrium
Evan Munro, Xu Kuang, and Stefan Wager. Treatment effects in market equilibrium. American Economic Review, forthcoming, 2025
2025
-
[64]
Optimal dynamic treatment regimes
Susan A Murphy. Optimal dynamic treatment regimes. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 65 0 (2): 0 331--355, 2003
2003
-
[65]
A generalization error for q-learning
Susan A Murphy. A generalization error for q-learning. Journal of Machine Learning Research, 6: 0 1073--1097, 2005
2005
-
[66]
Admission control in the presence of arrival forecasts with blocking-based policy optimization
Karthyek Murthy, Divya Padmanabhan, and Satyanath Bhat. Admission control in the presence of arrival forecasts with blocking-based policy optimization. In 2022 Winter Simulation Conference (WSC), pages 2270--2281. IEEE, 2022
2022
-
[67]
The regulation of queue size by levying tolls
Pinhas Naor. The regulation of queue size by levying tolls. Econometrica, 37 0 (1): 0 15--24, 1969
1969
-
[68]
Quasi-oracle estimation of heterogeneous treatment effects
Xinkun Nie and Stefan Wager. Quasi-oracle estimation of heterogeneous treatment effects. Biometrika, 108 0 (2): 0 299--319, 2021
2021
-
[69]
Learning when-to-treat policies
Xinkun Nie, Emma Brunskill, and Stefan Wager. Learning when-to-treat policies. Journal of the American Statistical Association, 116 0 (533): 0 392--409, 2021
2021
-
[70]
A Direct Search Optimization Method That Models the Objective and Constraint Functions by Linear Interpolation, pages 51--67
Michael JD Powell. A Direct Search Optimization Method That Models the Objective and Constraint Functions by Linear Interpolation, pages 51--67. Springer Netherlands, Dordrecht, 1994
1994
-
[71]
Direct search algorithms for optimization calculations
Michael JD Powell. Direct search algorithms for optimization calculations. Acta numerica, 7: 0 287--336, 1998
1998
-
[72]
Performance guarantees for individualized treatment rules
Min Qian and Susan A Murphy. Performance guarantees for individualized treatment rules. Annals of Statistics, 39 0 (2): 0 1180, 2011
2011
-
[73]
Optimal structural nested models for optimal sequential decisions
James M Robins. Optimal structural nested models for optimal sequential decisions. In Proceedings of the second seattle Symposium in Biostatistics, pages 189--326. Springer, 2004
2004
-
[74]
Introduction to Probability Models (Eleventh Edition)
Sheldon Ross. Introduction to Probability Models (Eleventh Edition). Academic Press, Boston, eleventh edition edition, 2014
2014
-
[75]
Average treatment effects in the presence of unknown interference
Fredrik S \"a vje, Peter M Aronow, and Michael G Hudgens. Average treatment effects in the presence of unknown interference. The Annals of Statistics, 49 0 (2): 0 673--701, 2021
2021
-
[76]
Treatment planning for victims with heterogeneous time sensitivities in mass casualty incidents
Yunting Shi, Nan Liu, and Guohua Wan. Treatment planning for victims with heterogeneous time sensitivities in mass casualty incidents. Operations Research, 72 0 (4): 0 1400--1420, 2024
2024
-
[77]
Hindsight learning for mdps with exogenous inputs
Sean R Sinclair, Felipe Vieira Frujeri, Ching-An Cheng, Luke Marshall, Hugo De Oliveira Barbalho, Jingling Li, Jennifer Neville, Ishai Menache, and Adith Swaminathan. Hindsight learning for mdps with exogenous inputs. In International Conference on Machine Learning, pages 3187...
2023
-
[78]
Minimax regret treatment choice with finite samples
J \"o rg Stoye. Minimax regret treatment choice with finite samples. Journal of Econometrics, 151 0 (1): 0 70--81, 2009
2009
-
[79]
Minimax regret treatment choice with covariates or with limited validity of experiments
J \"o rg Stoye. Minimax regret treatment choice with covariates or with limited validity of experiments. Journal of Econometrics, 166 0 (1): 0 138--156, 2012
2012
-
[80]
Treatment allocation under uncertain costs
Hao Sun, Evan Munro, Georgy Kalashnov, Shuyang Du, and Stefan Wager. Treatment allocation under uncertain costs. arXiv preprint arXiv:2103.11066, 2021
2021 arXiv
-
[81]
Empirical welfare maximization with constraints
Liyang Sun. Empirical welfare maximization with constraints. arXiv preprint arXiv:2103.15298, 2, 2021
2021
-
[82]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. A Bradford Book, Cambridge, MA, USA, 2018. ISBN 0262039249
2018
-
[83]
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour. Policy gradient methods for reinforcement learning with function approximation. Advances in neural information processing systems, 12, 1999
1999
-
[84]
Qini curves for multi-armed treatment rules
Erik Sverdrup, Han Wu, Susan Athey, and Stefan Wager. Qini curves for multi-armed treatment rules. Journal of Computational and Graphical Statistics, pages 1--13, 2024
2024
-
[85]
On causal inference in the presence of interference
Eric J Tchetgen Tchetgen and Tyler J VanderWeele. On causal inference in the presence of interference. Statistical Methods in Medical Research, 21 0 (1): 0 55--75, 2012
2012
-
[86]
Statistical treatment choice based on asymmetric minimax regret criteria
Aleksey Tetenov. Statistical treatment choice based on asymmetric minimax regret criteria. Journal of Econometrics, 166 0 (1): 0 157--165, 2012. doi:https://doi.org/10.1016/j.jeconom.2011.06.013
2012 doi
-
[87]
A. W. van der Vaart. Asymptotic Statistics. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 1998
1998
-
[88]
Learning and Value Function Approximation in Complex Decision Processes
Benjamin Van Roy. Learning and Value Function Approximation in Complex Decision Processes. PhD thesis, Massachusetts Institute of Technology, 1998
1998
-
[89]
Policy targeting under network interference
Davide Viviano. Policy targeting under network interference. Review of Economic Studies, page rdae041, 2024
2024
-
[90]
Estimation and inference of heterogeneous treatment effects using random forests
Stefan Wager and Susan Athey. Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association, 113 0 (523): 0 1228--1242, 2018
2018
-
[91]
Wainwright
Martin J. Wainwright. High-Dimensional Statistics: A Non-Asymptotic Viewpoint. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 2019
2019
-
[92]
Exploiting exogenous structure for sample-efficient reinforcement learning
Jia Wan, Sean R Sinclair, Devavrat Shah, and Martin J Wainwright. Exploiting exogenous structure for sample-efficient reinforcement learning. arXiv preprint arXiv:2409.14557, 2024
2024 arXiv
-
[93]
Learning optimal personalized treatment rules in consideration of benefit and risk: with an application to treating type 2 diabetes patients with insulin therapies
Yuanjia Wang, Haoda Fu, and Donglin Zeng. Learning optimal personalized treatment rules in consideration of benefit and risk: with an application to treating type 2 diabetes patients with insulin therapies. Journal of the American Statistical Association, 113 0 (521): 0 1--13, 2018
2018
-
[94]
Using future information to reduce waiting times in the emergency department via diversion
Kuang Xu and Carri W Chan. Using future information to reduce waiting times in the emergency department via diversion. Manufacturing & Service Operations Management, 18 0 (3): 0 314--331, 2016
2016
-
[95]
Estimating the optimal individualized treatment rule from a cost-effectiveness perspective
Yizhe Xu, Tom H Greene, Adam P Bress, Brian C Sauer, Brandon K Bellows, Yue Zhang, William S Weintraub, Andrew E Moran, and Jincheng Shen. Estimating the optimal individualized treatment rule from a cost-effectiveness perspective. Biometrics, 78 0 (1): 0 337--351, 2022
2022
-
[96]
Evaluating treatment prioritization rules via rank-weighted average treatment effects
Steve Yadlowsky, Scott Fleming, Nigam Shah, Emma Brunskill, and Stefan Wager. Evaluating treatment prioritization rules via rank-weighted average treatment effects. Journal of the American Statistical Association, 120 0 (549): 0 38--51, 2025
2025
-
[97]
Bossarte, Sarah M
Nur Hani Zainal, Robert M. Bossarte, Sarah M. Gildea, Irving Hwang, Chris J. Kennedy, Howard Liu, Alex Luedtke, Brian P. Marx, Maria V. Petukhova, Edward P. Post, Eric L. Ross, Nancy A. Sampson, Erik Sverdrup, Brett Turner, Stefan Wager, and Ronald C. Kessler. Developing an in...
2024
-
[98]
Estimating treatment effects under recommender interference: A structured neural networks approach
Ruohan Zhan, Shichao Han, Yuchen Hu, and Zhenling Jiang. Estimating treatment effects under recommender interference: A structured neural networks approach. arXiv preprint arXiv:2406.14380, 2024
2024
-
[99]
Robust estimation of optimal dynamic treatment regimes for sequential treatment decisions
Baqun Zhang, Anastasios A Tsiatis, Eric B Laber, and Marie Davidian. Robust estimation of optimal dynamic treatment regimes for sequential treatment decisions. Biometrika, 100 0 (3): 0 681--694, 2013
2013
-
[100]
Individualized policy evaluation and learning under clustered network interference
Yi Zhang and Kosuke Imai. Individualized policy evaluation and learning under clustered network interference. arXiv preprint arXiv:2311.02467, 2023
2023 arXiv
-
[101]
On-policy deep reinforcement learning for the average-reward criterion
Yiming Zhang and Keith W Ross. On-policy deep reinforcement learning for the average-reward criterion. In International Conference on Machine Learning, pages 12535--12545. PMLR, 2021
2021
-
[102]
Offline multi-action policy learning: Generalization and optimization
Zhengyuan Zhou, Susan Athey, and Stefan Wager. Offline multi-action policy learning: Generalization and optimization. Operations Research, 71 0 (1): 0 148--183, 2023
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.