REVIEW 2 major objections 7 minor 133 references
On Fisher Consistency of Surrogate Losses for Optimal Dynamic Treatment Regimes with Multiple Categorical Treatments per Stage
T0 review · 2 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Within the class of non-negative, stagewise separable surrogate losses, Fisher consistency for optimal dynamic treatment regimes holds exactly when each single-stage component satisfies Condition N1 and every component at stages t ≥ 2…
desk verdict A substantial sufficiency-side contribution, but the paper's headline necessary-and-sufficient claim is false as stated for P0: a constant shift of their own surrogates refutes necessity. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the single-stage functional Ψ*_t(p) = sup_{x∈$R^{{k_t}}$} Σ_{a∈[k_t]} p_a φ_t(x;a), the support function of the surrogate's image set V_φ = {φ(x;·) : x∈R^k}. Condition N1 states that maximizers of Ψ_t(·;p) always have pred(x) in the argmax of p; Condition N2 states that Ψ*_t(p) = C_φ max(p), which by Lemma 4.2 is equivalent to the scaled simplex C_φ $S^{{k-1}}$ lying inside conv(V_φ). The proof of the main theorems shows that when N1 and N2 hold, the surrogate regret V^ψ_* − V^ψ(f) is bounded below by a linear multiple of the true regret V* − V(f), and conversely constructs distributions showing that when these conditions fail, a maximizing score sequence can approach the surrogate optimum while the induced policies miss the optimal treatment regime.
What would settle it
Construct a product-form surrogate whose first-stage component satisfies Condition N1 and whose later-stage component satisfies N1 but not N2 (for instance, a shifted surrogate from Lemma 4.3), then search for a distribution satisfying Assumptions I–V with Y_t > C_min > 0 under which maximizing the surrogate value still yields an optimal DTR; finding such a pair would refute the necessity direction for P0, and a smooth concave PERM loss that is Fisher consistent for T = 2 would refute Theorem 3.1.
Extended reading notes
Core claim
The central discovery is an exact necessary-and-sufficient condition for Fisher consistency of separable surrogates in general multi-stage DTR problems. A surrogate of the product form ψ = ∏_t φ_t is Fisher consistent with respect to all distributions satisfying Assumptions I–V if and only if each single-stage φ_t satisfies Condition N1 (the natural single-stage calibration condition: a score vector that does not point to an optimal treatment cannot maximize the surrogate risk) and, for every t ≥ 2, φ_t additionally satisfies Condition N2, meaning its sup-functional Ψ*_t(p) = sup_x Σ_i p_i φ_t(x;i) equals C_φ max(p). Geometrically, N2 forces the closed convex hull of the image set {φ(x; ·)} to contain a scaled probability simplex, the image hull of the original discontinuous loss. The paper proves sufficiency for all T and necessity under non-negative rewards (Assumptions I–IV with Y_t ≥ 0), and argues the same conditions should hold for the strictly positive reward class P0.
Load-bearing premise
The necessity half of the main theorem is proved only for distributions whose rewards are non-negative (Y_t ≥ 0), and the paper asserts the same conditions are necessary for the strictly-positive-reward class P0 without proving that transfer; if a separable surrogate were Fisher consistent on P0 while violating Condition N2, the 'if and only if' claim as stated would fail.
Editorial extensions
If this is right
- Maximizing any product-form surrogate that satisfies Conditions N1 and N2 recovers an optimal dynamic treatment regime exactly, while violating either condition allows a maximizing sequence to converge to a suboptimal policy.
- No smooth concave permutation-equivariant and relative-margin-based (PERM) surrogate can be Fisher consistent for two or more stages, so convexification of simultaneous direct search is impossible within that class.
- The proposed kernel-based and product-based surrogates satisfy N1–N3, yielding smooth non-concave losses with a linear regret constant and a regret decay rate n^{-(1+α)/(2+α+q/β)} up to log factors under small-noise and Hölder-smooth blip assumptions.
- For restricted policy classes such as linear policies, Conditions N1–N3 still deliver regret bounds when the optimal policy lies inside the class, giving a formal guarantee for interpretable SDSS implementations.
- Shifted single-stage surrogates that satisfy Condition N1 but fail Condition N2 are provably not Fisher consistent for T ≥ 2, so the multi-stage setting demands surrogates geometrically closer to the original discontinuous loss.
Reading between the lines
- The characterization points to a 'no free lunch' geometry: any Fisher consistent surrogate for multi-stage DTR must have an image hull large enough to contain a scaled simplex, a strictly stronger requirement than single-stage calibration, which explains why the multi-stage problem resists convexification.
- The necessity-direction proof gap (Y_t ≥ 0 versus Y_t > C_min > 0) might be patchable by a location-shift argument, but the paper does not supply it; a counterexample on the strictly positive class P0 would reveal a genuine boundary phenomenon where the iff statement fails as written.
- The SDSS regret analysis treats optimization error as an additive black box, yet the paper's own toy example shows gradient methods can get trapped in suboptimal plateaus, so practical performance depends critically on the heuristics (restarts, minibatching, learning-rate scheduling) rather than on the surrogate conditions alone.
- The product-based surrogate has a slower approximation error (δ^α) than the symmetric kernel-based surrogate (δ^{1+α}), so among Fisher consistent surrogates the choice affects the achievable rate—an implicit trade-off the paper notes but leaves for future selection.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies the sequential DTR classification problem with T stages and k_t ≥ 2 treatment options per stage under known propensity scores, and it contributes both theory and methodology. The central theoretical claim is a characterization of Fisher consistency within the class of non-negative stagewise separable surrogates ψ(x₁,…,x_T; a₁,…,a_T) = ∏_{t=1}^T φ_t(x_t; a_t): Theorem 4.1 shows that Condition N1 for every stage plus Condition N2 for every stage t ≥ 2 are sufficient for Fisher consistency on the primary class P₀ (Assumptions I–V with Y_t > C_min > 0), and Theorem 4.2 shows that F1–F2 are necessary for Fisher consistency on the larger class of distributions with Y_t ≥ 0; the paper then refers to F1–F2 as the necessary and sufficient conditions for the separable class. Additional contributions include an impossibility theorem for smooth concave PERM surrogates (Theorem 3.1), explicit smooth non-convex surrogates (kernel- and product-based) with closed-form constants, the SDSS optimization algorithm, a regret bound under a Tsybakov small-noise condition (Theorem 6.2, Corollary 1), and an empirical evaluation on simulations and on MIMIC-III sepsis data.
Significance. The paper's strengths are substantial. The sufficient conditions and surrogate constructions are concrete, closed-form, and checkable (explicit χ_φ, J_t, and C* constants), which makes the N1–N2 conditions readily falsifiable against new surrogate proposals; the PERM impossibility result is non-trivial and goes beyond prior binary-treatment work; and the regret analysis attains a rate matching the binary-classification benchmark under comparable Tsybakov-type conditions. The empirical study is careful, with publicly available code, standard comparators (Q-learning and ACWL), and a sensible AIPW-based evaluation on MIMIC-III EHR data. If the headline characterization were fully established for P₀, this would be the first necessary-and-sufficient DTR Fisher-consistency result in a non-trivial surrogate class. The main weakness is that the necessity half of the characterization is proved only for the larger Y_t ≥ 0 class, not for the paper's primary class P₀, while the Abstract and Introduction present the result as a complete necessary-and-sufficient characterization; the body of the paper acknowledges this gap explicitly but then sets it aside.
major comments (2)
- [Section 4, Theorems 4.1–4.2; Abstract] The headline claim of a necessary-and-sufficient characterization for the paper's primary class P₀ is not established by the theorems as stated. Theorem 4.2 proves necessity of F1–F2 under the hypothesis that ψ is Fisher consistent on the strictly larger class of distributions satisfying Assumptions I–IV with Y_t ≥ 0. Because P₀ (Assumption V: Y_t > C_min > 0) is a proper subset of that class, Fisher consistency on P₀ is a weaker property than Fisher consistency on the larger class; the implication 'FC on {Y_t ≥ 0} implies F1+F2' therefore does not transfer downward to P₀. The paper's own text acknowledges this ('It may be possible to show that F1 and F2 are necessary for Fisher consistency even under P₀... considerably more technical') but then states 'Despite this minor discrepancy, we will refer to F1 and F2 as the necessary and sufficient conditions...'. The discrepancy is not minor for the central claim: the necessity of F1+F2 for P₀ remains an open proof obligation, and the Abstract's claim to establish 'necessary and sufficient conditions for DTR Fisher consistency within the class of non-negative, stagewise separable surrogate losses' overstates what Theorems 4.1–4.2 prove. The authors should either supply the P₀-necessity proof or visibly restrict the characterization to the Y_t ≥ 0 class and present the P₀ direction as a conjecture.
- [Section 4, Lemma 4.3; Section 2, Definition 2.1] A natural attempt to falsify the P₀-necessity claim—adding a positive constant to a surrogate satisfying N1+N2, as permitted by Lemma 4.3, and noting that the shift violates N2—does not produce a counterexample, because the shifted surrogate is not Fisher consistent on all of P₀. Concretely, take T = 2, k₁ = k₂ = 2, no covariates, π₁ = π₂ = 1/2, deterministic outcomes with E[Y₁+Y₂ | A₁,A₂] = [[10, 9], [1, 10.1]] (realizable with Y₁ ≡ 0.1 and Y₂ ∈ {9.9, 8.9, 0.9, 10.0}, so the distribution lies in P₀), and the logistic shifted surrogate φ(x; i) = σ(x_i − x_{3−i}) + γ with γ = 0.1. Writing s = σ(f₁₁ − f₁₂), u = σ(f₂(1)₁ − f₂(1)₂), v = σ(f₂(2)₁ − f₂(2)₂), direct computation gives V_ψ(s,u,v) = (s + 0.1)(u + 10.9) + (1.1 − s)(11.21 − 9.1v). The supremum, 14.211, is approached only as s → 1, u → 1, v → 0, so every maximizing sequence has pred(f₁) = 1 and pred(f₂(1)) = 1; consequently V(f_m) → W(1,1) = 10, while V* = 10.1 is attained only by d₁ = 2. The shifted surrogate thus fails Fisher consistency at this P ∈ P₀. The proof gap in the necessity direction for P₀ stands, but the available evidence points to the characterization being true rather than false; the manuscript should close the gap with a proof or state the P₀-necessity as an open problem.
minor comments (7)
- [Section 3, Example 1] The sentence after Result 1, 'the loss in (3.1) can not be Fisher inconsistent', states the opposite of the intended conclusion and should read 'cannot be Fisher consistent'.
- [Section 1.1.1] The heading 'Thoretical results on Fisher consistency' contains a typo; it should read 'Theoretical results on Fisher consistency'.
- [Equation (5.4)] Equation (5.4) contains the stray symbol '≡='; a single equals sign is evidently intended.
- [Table 1] The table conflates the score-vector columns with the argmax-set column; since the pred link returns max(argmax), the all-zero score vectors in Settings 1–2 yield pred = 3, and this tie-breaking should be stated explicitly where the table is discussed.
- [Section 4.3, Lemma 4.8] In the product-based template, the i = 1 case is written as ∏_{j∈[k]} τ(y_j), but (4.11) and the definition Δx₁ = (x₁ − x₂, …, x₁ − x_k) require the product to run over j ∈ [k] \ {1}, excluding the own-margin term; the same indexing issue affects (4.14).
- [Section 4.0.2] Immediately after Theorem 4.1, the sentence 'The proof of Theorem 4.2 can be found in Supplement S7' must refer to Theorem 4.1, since the following sentence describes the sufficiency construction.
- [Section 2, Assumption V] The 'without loss of generality' location-shift justification is valid for the value function because V'(d) = V(d) + T·c preserves the argmax, but the text should state explicitly that the surrogate analysis is applied to the shifted outcomes and that the shift changes the finite-sample surrogate objective, a point the paper itself later raises in Section 9.
Circularity Check
No significant circularity. The necessary-and-sufficient characterization is derived in the paper from the definitions and from standard external calibration results; self-citations are contextual rather than load-bearing. An acknowledged proof gap in the necessity direction (Y_t >= 0 vs P0) is a correctness concern, not a circularity.
full rationale
The paper's central claim is a characterization theorem: within the class of non-negative stagewise separable surrogates, DTR Fisher consistency holds if and only if each single-stage component satisfies Condition N1 and, for t >= 2, Condition N2. The proof derives these conditions from the definition of Fisher consistency via constructed 'good sets' of distributions, rather than assuming the conditions or fitting them to the target. No numeric parameters are fitted and then renamed as predictions; the kernel-based and product-based surrogates are defined in closed form and then verified to satisfy the stated conditions. The necessity proof in Theorem 4.2 is conducted for the broader class of distributions with Y_t >= 0, and the paper explicitly acknowledges that transferring necessity to P0 (Y_t > C_min > 0) is not proved: 'It may be possible to show that F1 and F2 are necessary for Fisher consistency even under P0... the corresponding proof would be considerably more technical.' This is a proof gap and a possible correctness defect, but it is not circularity: it does not reduce the conclusion to an input by construction. Self-citations to Laha et al. (2024a) are used for comparison, for examples, and for prior binary-treatment results; the main separable-surrogate characterization is proved in this paper from the definitions and from standard external multiclass calibration results (Zhang 2004; Tewari and Bartlett 2007; Wang and Scott 2023a,b), which are independent support. The Theorem 4.1 sufficient direction and the regret bounds are likewise derived rather than assumed. No equation in the paper is equivalent to its own input by construction, so the appropriate circularity score is 1.
Assumptions & free parameters
assumptions (6)
- domain assumption Identifiability assumptions I-III: positivity, consistency, and sequential ignorability for the IPW representation of the value function (Equation 2.2).
- domain assumption Assumptions IV-V: outcomes are bounded above and can be shifted so that Y_t > C_min > 0 for all t.
- domain assumption Tsybakov small noise condition (Assumption 1) on the gap in optimal Q-functions.
- domain assumption Hölder smoothness of blip functions (Assumption 2) for the neural network rate in Corollary 1.
- standard math Known multiclass Fisher consistency characterizations (Tewari and Bartlett 2007, Zhang 2004, Wang and Scott 2023a,b).
- standard math Convex analysis facts: support functions, closed convex hulls, proper and closed convex functions, and strict convexity conditions on the negative template.
Cite this review
Pith. "Pith review of On Fisher Consistency of Surrogate Losses for Optimal Dynamic Treatment Regimes with Multiple Categorical Treatments per Stage." pith.science (2026). https://pith.science/paper/BLOUEL5Q
@misc{pith2026250517285,
author = {Pith},
title = {Pith review of: On Fisher Consistency of Surrogate Losses for Optimal Dynamic Treatment Regimes with Multiple Categorical Treatments per Stage},
year = {2026},
howpublished = {\url{https://pith.science/paper/BLOUEL5Q}},
note = {Machine review of arXiv:2505.17285}
}
read the original abstract
Patients with chronic diseases often receive treatments at multiple time points, or stages. Our goal is to learn the optimal dynamic treatment regime (DTR) from longitudinal patient data. When both the number of stages and the number of treatment levels per stage are arbitrary, estimating the optimal DTR reduces to a sequential, weighted, multiclass classification problem (Kosorok and Laber, 2019). In this paper, we aim to solve this classification problem simultaneously across all stages using Fisher consistent surrogate losses. Although computationally feasible Fisher consistent surrogates exist in special cases, e.g., the binary treatment setting, a unified theory of Fisher consistency remains largely unexplored. We establish necessary and sufficient conditions for DTR Fisher consistency within the class of non-negative, stagewise separable surrogate losses. To our knowledge, this is the first result in the DTR literature to provide necessary conditions for Fisher consistency within a non-trivial surrogate class. Furthermore, we show that many convex surrogate losses fail to be Fisher consistent for the DTR classification problem, and we formally establish this inconsistency for smooth, permutation equivariant, and relative-margin-based convex losses. Building on this, we propose SDSS (Simultaneous Direct Search with Surrogates), which uses smooth, non-concave surrogate losses to learn the optimal DTR. We develop a computationally efficient, gradient-based algorithm for SDSS. When the optimization error is small, we establish a sharp upper bound on SDSS's regret decay rate. We evaluate the numerical performance of SDSS through simulations and demonstrate its real-world applicability by estimating optimal fluid resuscitation strategies for severe septic patients using electronic health record data.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Akhtar, M., Tanveer, M., and Arshad, M. (2024). Hawkeye: advancing robust regression with bounded, smooth, and insensitive loss function. arXiv preprint arXiv:2401.16785\/
work page Pith review arXiv 2024
-
[2]
and Wager, S
Athey, S. and Wager, S. (2021). Policy learning with observational data. Econometrica\/ , 89 (1), 133--161
2021
-
[3]
and Tsybakov, A
Audibert, J.-Y. and Tsybakov, A. B. (2007). Fast learning rates for plug-in classifiers. The Annals of statistics\/ , 35 (2), 608--633
2007
-
[4]
L., Jordan, M
Bartlett, P. L., Jordan, M. I., and McAuliffe, J. D. (2006). Convexity, classification, and risk bounds. Journal of the American Statistical Association\/ , 101 (473), 138--156
2006
-
[5]
Beijbom, O., Saberian, M., Kriegman, D., and Vasconcelos, N. (2014). Guess-averse loss functions for cost-sensitive multiclass boosting. In International Conference on Machine Learning\/ , pages 586--594. PMLR
2014
-
[6]
and Kallus, N
Bennett, A. and Kallus, N. (2020). Efficient policy learning from surrogate-loss classification reductions. In International Conference on Machine Learning\/ , pages 788--798. PMLR
2020
-
[7]
Bhattacharyya, A., Ghoshal, S., and Saket, R. (2018). Hardness of learning noisy halfspaces using polynomial thresholds. In Conference On Learning Theory\/ , pages 876--917. PMLR
2018
-
[8]
Billingsley, P. (2017). Probability and measure\/ . John Wiley & Sons
2017
Show all 133 references
-
[9]
E., and Nocedal, J
Bottou, L., Curtis, F. E., and Nocedal, J. (2018). Optimization methods for large-scale machine learning. SIAM review\/ , 60 (2), 223--311
2018
-
[10]
and Vandenberghe, L
Boyd, S. and Vandenberghe, L. (2004). Convex Optimization\/ . Cambridge University Press, New York, NY, USA
2004
-
[11]
Calauzenes, C., Usunier, N., and Gallinari, P. (2012). On the (non-) existence of convex, calibrated surrogate losses for ranking. Advances in Neural Information Processing Systems 25 (NIPS 2012)\/ , pages 197--205
2012
-
[12]
and Moodie, E
Chakraborty, B. and Moodie, E. E. (2013). Statistical methods for dynamic treatment regimes\/ , volume 2. Springer
2013
-
[13]
Chapagain, N. (2024a). Data application code for SDSS method. https://github.com/nilson01/ApplicationFin . GitHub repository
2024
-
[14]
Chapagain, N. (2024b). Simulation code for SDSS method. https://github.com/nilson01/SimulationDirectSearch. GitHub repository
2024
-
[15]
Charoenphakdee, N., Lee, J., and Sugiyama, M. (2019). On symmetric losses for learning from corrupted labels. In International Conference on Machine Learning\/ , pages 961--970. PMLR
2019
-
[16]
R., and Liu, Y
Chen, J., Fu, H., He, X., Kosorok, M. R., and Liu, Y. (2018). Estimating individualized treatment rules for ordinal treatments. Biometrics\/ , 74 (3), 924--933
2018
-
[17]
Chen, Y., Chi, Y., Fan, J., and Ma, C. (2019). Gradient descent with random initialization: Fast global convergence for nonconvex phase retrieval. Mathematical Programming\/ , 176 , 5--37
2019
-
[18]
N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014). Identifying and attacking the saddle point problem in high-dimensional non-convex optimization. In Advances in Neural Information Processing Systems\/ , pages 2933--2941
2014
-
[19]
J., and Chakraborty, B
Deliu, N., Williams, J. J., and Chakraborty, B. (2022). Reinforcement learning in modern biostatistics: constructing optimal adaptive interventions. arXiv preprint arXiv:2203.02605\/
2022 arXiv
-
[20]
N., Dao, C
Do, S. N., Dao, C. X., Nguyen, T. A., Nguyen, M. H., Pham, D. T., Nguyen, N. T., Huynh, D. Q., Hoang, Q. T. A., Van Bui, C., Vu, T. D., et al. (2023). Sequential organ failure assessment (sofa) score for predicting mortality in patients with sepsis in vietnamese intensive care...
2023
-
[21]
Doss, C. R. and Wellner, J. A. (2016). Global rates of convergence of the MLEs of log-concave and s -concave densities. Ann. Statist. , 44 , 954--981
2016
-
[22]
Doss, C. R. and Wellner, J. A. (2019). Univariate log-concave density estimation with symmetry or modal constraints. Electron. J. Stat. , 13 , 2391--2461
2019
-
[23]
C., Mackey, L
Duchi, J. C., Mackey, L. W., and Jordan, M. I. (2010). On the consistency of ranking algorithms. In ICML\/
2010
-
[24]
E., and Telgarsky, M
Dud \' k, M., Schapire, R. E., and Telgarsky, M. (2022). Convex analysis at infinity: An introduction to astral space. arXiv preprint arXiv:2205.03260\/
2022
-
[25]
Fathony, R., Liu, A., Asif, K., and Ziebart, B. (2016). Adversarial multiclass classification: A risk minimization perspective. Advances in Neural Information Processing Systems\/ , 29
2016
-
[26]
Feng, H., Ning, Y., and Zhao, J. (2022). Nonregular and minimax estimation of individualized thresholds in high dimension with binary responses. The Annals of Statistics\/ , 50 (4), 2284--2305
2022
-
[27]
Finocchiaro, J., Frongillo, R., and Waggoner, B. (2019). An embedding framework for consistent polyhedral surrogates. Advances in neural information processing systems\/ , 32
2019
-
[28]
and Zhou, Z.-H
Gao, W. and Zhou, Z.-H. (2011). On the consistency of multi-label learning. In Proceedings of the 24th annual conference on learning theory\/ , pages 341--358
2011
-
[29]
and Nickl, R
Giné, E. and Nickl, R. (2015). Mathematical Foundations of Infinite-Dimensional Statistical Models\/ , volume 40 of Cambridge series in statistical and probabilistic mathematics\/ . Cambridge University Press, Cambridge
2015
-
[30]
Glasmachers, T., Igel, C., et al. (2016). A unified view on multi-class support vector classification. Journal of Machine Learning Research\/ , 17 (45), 1--32
2016
-
[31]
and Bengio, Y
Glorot, X. and Bengio, Y. (2010). Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics\/ , pages 249--256. JMLR Workshop and Conference Proceedings
2010
-
[32]
Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep Learning\/ . MIT Press. http://www.deeplearningbook.org
2016
-
[33]
A., and Davidian, M
Hager, R., Tsiatis, A. A., and Davidian, M. (2018). Optimal two-stage dynamic treatment regimes from a classification perspective with censored survival data. Biometrics\/ , 74 (4), 1180--1192
2018
-
[34]
Han, S. (2021). Identification in nonparametric models for dynamic treatment effects. Journal of Econometrics\/ , 225 (2), 132--147
2021
-
[35]
He, K., Zhang, X., Ren, S., and Sun, J. (2015). Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE International Conference on Computer Vision\/ , pages 1026--1034
2015
-
[36]
W., and Ridder, G
Hirano, K., Imbens, G. W., and Ridder, G. (2003). Efficient estimation of average treatment effects using the estimated propensity score. Econometrica\/ , 71 (4), 1161--1189
2003
-
[37]
and Lemar \'e chal, C
Hiriart-Urruty, J.-B. and Lemar \'e chal, C. (2004). Fundamentals of convex analysis\/ . Springer Science & Business Media
2004
-
[38]
Hochreiter, S., Bengio, Y., Frasconi, P., Schmidhuber, J., et al. (2001). Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
2001
-
[39]
Horn, R. A. and Johnson, C. R. (2012). Matrix analysis\/ . Cambridge university press
2012
-
[40]
Horowitz, J. L. (1992). A smoothed maximum score estimator for the binary response model. Econometrica: journal of the Econometric Society\/ , pages 505--531
1992
-
[41]
Jiang, B., Song, R., Li, J., and Zeng, D. (2019). Entropy learning for dynamic treatment regimes. Statistica Sinica\/ , 29 (4), 1633
2019
-
[42]
and Li, L
Jiang, N. and Li, L. (2016). Doubly robust off-policy value evaluation for reinforcement learning. In International Conference on Machine Learning\/ , pages 652--661. PMLR
2016
-
[43]
Jin, C., Liu, Q., and Miryoosefi, S. (2021). Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms. Advances in neural information processing systems\/ , 34 , 13406--13418
2021
-
[44]
and Pollard, T
Johnson, A. and Pollard, T. (2018). sepsis3-mimic
2018
-
[45]
Johnson, A., Pollard, T., Shen, L., Lehman, L., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Celi, L., and Mark, R. (2016). Mimic-iii, a freely accessible critical care database. Scientific Data\/ , 3 , 160035
2016
-
[46]
E., Aboab, J., Raffa, J
Johnson, A. E., Aboab, J., Raffa, J. D., Pollard, T. J., Deliberato, R. O., Celi, L. A., and Stone, D. J. (2018). A comparative analysis of sepsis identification methods in an electronic database. Critical care medicine\/ , 46 (4), 494
2018
-
[47]
and Uehara, M
Kallus, N. and Uehara, M. (2020). Double reinforcement learning for efficient off-policy evaluation in markov decision processes. Journal of Machine Learning Research\/ , 21 (167)
2020
-
[48]
Kingma, D. P. and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980\/
2014 arXiv
-
[49]
Koltchinskii, V. (2009). 2008 saint flour lectures oracle inequalities in empirical risk minimization and sparse recovery problems
2009
-
[50]
Komorowski, M., Celi, L., Badawi, O., Gordon, A., and Faisal, A. (2018). The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care. Nature Medicine\/ , 24 (11), 1716--1720
2018
-
[51]
Kosorok, M. R. and Laber, E. B. (2019). Precision medicine. Annual Review of Statistics and its Application \/ , 6 , 263--286
2019
-
[52]
B., Lizotte, D
Laber, E. B., Lizotte, D. J., Qian, M., Pelham, W. E., and Murphy, S. A. (2014). Dynamic treatment regimes: Technical challenges and applications. Electronic journal of statistics\/ , 8 (1), 1225
2014
-
[53]
Laha, N. (2021). Adaptive estimation in symmetric location model under log-concavity constraint. Electronic Journal of Statistics \/ , 15 , 2939--3014
2021
-
[54]
Laha, N., Sonabend-W., A., Mukherjee, R., and Cai, T. (2024a). Finding the optimal dynamic treatment regimes using smooth fisher consistent surrogate loss. Annals of Statistics\/ , 52 (2), 679--707
2024
-
[55]
Laha, N., Sonabend-W, A., Mukherjee, R., and Cai, T. (2024b). Finding the optimal dynamic treatment regimes using smooth fisher consistent surrogate loss
2024
-
[56]
M., and De Backer, D
Lat, I., Coopersmith, C. M., and De Backer, D. (2021). The surviving sepsis campaign: fluid resuscitation and vasopressor therapy research priorities in adult patients. Intensive Care Medicine Experimental\/ , 9 (1), 1--16
2021
-
[57]
Lee, Y., Lin, Y., and Wahba, G. (2004). Multicategory support vector machines: Theory and application to the classification of microarray data and satellite radiance data. Journal of the American Statistical Association\/ , 99 (465), 67--81
2004
-
[58]
Liu, M., Shen, X., and Pan, W. (2021). Outcome weighted -learning for individualized treatment rules. Stat\/ , 10 (1), e343
2021
-
[59]
Liu, M., Wang, Y., Fu, H., and Zeng, D. (2024a). Controlling cumulative adverse risk in learning optimal dynamic treatment regimens. Journal of the American Statistical Association\/ , 119 (548), 2622--2633
2024
-
[60]
Liu, M., Wang, Y., Fu, H., and Zeng, D. (2024b). Learning optimal dynamic treatment regimens subject to stagewise risk controls. Journal of Machine Learning Research\/ , 25 (128), 1--64
2024
-
[61]
Liu, Y. (2007). Fisher consistency of multicategory support vector machines. In Artificial intelligence and statistics\/ , pages 291--298
2007
-
[62]
and Shen, X
Liu, Y. and Shen, X. (2006). Multicategory -learning. Journal of the American Statistical Association\/ , 101 (474), 500--509
2006
-
[63]
R., Zhao, Y., and Zeng, D
Liu, Y., Wang, Y., Kosorok, M. R., Zhao, Y., and Zeng, D. (2018). Augmented outcome-weighted learning for estimating optimal dynamic treatment regimens. Statistics in medicine\/ , 37 (26), 3776--3788
2018
-
[64]
Luedtke, A. R. and Van Der Laan, M. J. (2016). Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy. Annals of statistics\/ , 44 (2), 713
2016
-
[65]
Marik, P. (2015). The demise of early goal-directed therapy for severe sepsis and septic shock. Acta Anaesthesiologica Scandinavica\/ , 59 (5), 561--567
2015
-
[66]
and Vasconcelos, N
Masnadi-Shirazi, H. and Vasconcelos, N. (2008). On the design of loss functions for classification: theory, robustness to outliers, and savageboost. Advances in neural information processing systems\/ , 21
2008
-
[67]
Meng, H., Zhao, Y.-Q., Fu, H., and Qiao, X. (2020). Near-optimal individualized treatment recommendations. J. Mach. Learn. Res. , 21 , 183--1
2020
-
[68]
E., Richardson, T
Moodie, E. E., Richardson, T. S., and Stephens, D. A. (2007). Demystifying optimal dynamic treatment regimes. Biometrics\/ , 63 (2), 447--455
2007
-
[69]
E., Dean, N., and Sun, Y
Moodie, E. E., Dean, N., and Sun, Y. R. (2014). Q-learning: Flexible learning about useful utilities. Statistics in Biosciences\/ , 6 , 223--243
2014
-
[70]
Murphy, S. A. (2003). Optimal dynamic treatment regimes. Journal of the Royal Statistical Society: Series B (Statistical Methodology)\/ , 65 (2), 331--355
2003
-
[71]
Murphy, S. A. (2005). A generalization error for Q-learning . Journal Of Machine Learning Research\/ , 6 , 1073--1097
2005
-
[72]
A., van der Laan, M
Murphy, S. A., van der Laan, M. J., and Robins, J. M. (2001a). Marginal mean models for dynamic regimes. Journal of the American Statistical Association\/ , 96 (456), 1410--1423
2001
-
[73]
A., van der Laan, M
Murphy, S. A., van der Laan, M. J., and Robins, J. M. (2001b). Marginal mean models for dynamic regimes. Journal of the American Statistical Association\/ , 96 (456), 1410--1423
2001
-
[74]
S., Ravikumar, P
Natarajan, N., Dhillon, I. S., Ravikumar, P. K., and Tewari, A. (2013). Learning with noisy labels. Advances in neural information processing systems\/ , 26
2013
-
[75]
S., and Cai, T
Neykov, M., Liu, J. S., and Cai, T. (2016). On the characterization of a class of Fisher-consistent loss functions and its application to boosting. The Journal of Machine Learning Research\/ , 17 (1), 2498--2529
2016
-
[76]
and Sanner, S
Nguyen, T. and Sanner, S. (2013). Algorithms for direct 0--1 loss optimization in binary classification. In International conference on machine learning\/ , pages 1085--1093. PMLR
2013
-
[77]
Orellana, L., Rotnitzky, A., and Robins, J. M. (2010). Dynamic regime marginal structural mean models for estimation of optimal dynamic treatment regimes, part i: main content. The international journal of biostatistics\/ , 6 (2)
2010
-
[78]
Osokin, A., Bach, F., and Lacoste-Julien, S. (2017). On structured prediction theory with calibrated convex surrogate losses. Advances in Neural Information Processing Systems\/ , 30
2017
-
[79]
and Zhao, Y.-Q
Pan, Y. and Zhao, Y.-Q. (2021). Improved doubly robust estimation in learning optimal individualized treatment rules. Journal of the American Statistical Association\/ , 116 (533), 283--294
2021
-
[80]
Patrini, G., Rozza, A., Krishna Menon, A., Nock, R., and Qu, L. (2017). Making deep neural networks robust to label noise: A loss correction approach. In Proceedings of the IEEE conference on computer vision and pattern recognition\/ , pages 1944--1952
2017
-
[81]
and Murphy, S
Qian, M. and Murphy, S. A. (2011). Performance guarantees for individualized treatment rules. Annals of statistics\/ , 39 (2), 1180
2011
-
[82]
A., Szolovits, P., and Ghassemi, M
Raghu, A., Komorowski, M., Celi, L. A., Szolovits, P., and Ghassemi, M. (2017). Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach. abs/1705.08422
2017 arXiv
-
[83]
Ramaswamy, H. G. and Agarwal, S. (2016). Convex calibration dimension for multiclass loss matrices. The Journal of Machine Learning Research\/ , 17 (1), 397--441
2016
-
[84]
and Zeevi, A
Rigollet, P. and Zeevi, A. J. (2010). Nonparametric bandits with covariates. In Annual Conference Computational Learning Theory\/
2010
-
[85]
Robins, J. M. (1997). Causal inference from complex longitudinal data. In Latent variable modeling and applications to causality\/ , pages 69--117. Springer
1997
-
[86]
Robins, J. M. (2004). Optimal structural nested models for optimal sequential decisions. Proceedings of the second Seattle symposium on biostatistics In D. Y. Lin and P. Heagerty (Eds.) (pp. 189--326). New York: Springer
2004
-
[87]
M., Rotnitzky, A., and Zhao, L
Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association\/ , 89 (427), 846--866
1994
-
[88]
Rockafellar, R. T. (1970). Convex Analysis \/ . Princeton University Press
1970
-
[89]
Rosset, S., Zhu, J., and Hastie, T. (2003). Margin maximizing loss functions. Advances in neural information processing systems\/ , 16
2003
-
[90]
and Wellner, J
Saumard, A. and Wellner, J. A. (2014). Log-concavity and strong log-concavity: a review. Statistics surveys\/ , 8 , 45
2014
-
[91]
Nonparametric regression using deep neural networks with relu activation function
Schmidt-Hieber (2020). Nonparametric regression using deep neural networks with relu activation function. Annals of Statistics\/ , 48 (4), 1875--1897
2020
-
[92]
J., Tsiatis, A
Schulte, P. J., Tsiatis, A. A., Laber, E. B., and Davidian, M. (2014). Q - and A -learning methods for estimating optimal dynamic treatment regimes. Statist. Sci. , 29 (4), 640--661
2014
-
[93]
and Wellner, J
Seregin, A. and Wellner, J. A. (2010). Nonparametric estimation of multivariate convex-transformed densities. Annals of statistics\/ , 38 (6), 3751
2010
-
[94]
N., Cai, T., and Mukherjee, R
Sonabend-W, A., Laha, N., Ananthakrishnan, A. N., Cai, T., and Mukherjee, R. (2023). Semi-supervised off-policy reinforcement learning and value estimation for dynamic treatment regimes. Journal of Machine Learning Research\/ , 24 (323), 1--86
2023
-
[95]
Steinwart, I., Scovel, C., et al. (2007). Fast rates for support vector machines using gaussian kernels. The Annals of Statistics\/ , 35 (2), 575--607
2007
-
[96]
and Wang, L
Sun, Y. and Wang, L. (2021). Stochastic tree search for estimating optimal dynamic treatment regimes. Journal of the American Statistical Association\/ , 116 (533), 421--432
2021
-
[97]
Sutton, R. S. (2018). Reinforcement learning : an introduction\/ . Adaptive computation and machine learning. The MIT Press, Cambridge, Massachusetts ; London, England, second edition. edition
2018
-
[98]
and Wang, L
Tao, Y. and Wang, L. (2017). Adaptive contrast weighted learning for multi-stage multi-treatment decision-making. Biometrics\/ , 73 (1), 145--155
2017
-
[99]
Tao, Y., Wang, L., and Almirall, D. (2018). Tree-based reinforcement learning for estimating optimal dynamic treatment regimes. The annals of applied statistics\/ , 12 (3), 1914
2018
-
[100]
and Bartlett, P
Tewari, A. and Bartlett, P. L. (2007). On the consistency of multiclass classification methods. Journal of Machine Learning Research\/ , 8 (May), 1007--1025
2007
-
[101]
and Brunskill, E
Thomas, P. and Brunskill, E. (2016). Data-efficient off-policy policy evaluation for reinforcement learning. In International Conference on Machine Learning\/ , pages 2139--2148. PMLR
2016
-
[102]
Tsiatis, A. (2006). Semiparametric theory and missing data. 73
2006
-
[103]
A., Davidian, M., Holloway, S
Tsiatis, A. A., Davidian, M., Holloway, S. T., and Laber, E. B. (2019). Dynamic treatment regimes: Statistical methods for precision medicine\/ . Chapman and Hall/CRC
2019
-
[104]
Tsybakov, A. B. (2004). Optimal aggregation of classifiers in statistical learning. The Annals of Statistics\/ , 32 (1), 135--166
2004
-
[105]
and Lederer, J
Ven, L. and Lederer, J. (2021). Regularization and reparameterization avoid vanishing gradients in sigmoid-type networks. arXiv preprint arXiv:2106.02260\/
2021 arXiv
-
[106]
Wallace, M. P. and Moodie, E. E. (2015). Doubly-robust dynamic treatment regimen estimation via weighted least squares. Biometrics\/ , 71 (3), 636--644
2015
-
[107]
Wang, K., Pei, H., Cao, J., and Zhong, P. (2020). Robust regularized extreme learning machine for regression with non-convex loss function via dc program. Journal of the Franklin Institute\/ , 357 (11), 7069--7091
2020
-
[108]
and Scott, C
Wang, Y. and Scott, C. (2023a). On classification-calibration of gamma-phi losses. In The Thirty Sixth Annual Conference on Learning Theory\/ , pages 4929--4951. PMLR
2023
-
[109]
and Scott, C
Wang, Y. and Scott, C. (2023b). Unified binary and multiclass margin-based classification. arXiv preprint arXiv:2311.17778\/
2023 arXiv
-
[110]
Wang, Y., Fu, H., and Zeng, D. (2018). Learning optimal personalized treatment rules in consideration of benefit and risk: with an application to treating type 2 diabetes patients with insulin therapies. Journal of the American Statistical Association\/ , 113 (521), 1--13
2018
-
[111]
Watkins, C. J. C. H. (1989). Learning from Delayed Rewards\/ . Ph.D. thesis, King's College, Cambridge, UK
1989
-
[112]
Wellner, J. A. (2005). Empirical processes: Theory and applications. Lecture notes
2005
-
[113]
Weston, J., Watkins, C., et al. (1999). Support vector machines for multi-class pattern recognition. In Esann\/ , volume 99, pages 219--224
1999
-
[114]
and Liu, Y
Wu, Y. and Liu, Y. (2007). Robust truncated hinge loss support vector machines. Journal of the American Statistical Association\/ , 102 (479), 974--983
2007
-
[115]
Xu, T., Wang, J., and Fang, Y. (2014). A model-free estimation for the covariate-adjusted youden index and its associated cut-point. Statistics in medicine\/ , 33 (28), 4963--4974
2014
-
[116]
S., and Thall, P
Xu, Y., M \"u ller, P., Wahed, A. S., and Thall, P. F. (2016). Bayesian nonparametric estimation for dynamic treatment regimes with sequential transition times. Journal of the American Statistical Association\/ , 111 (515), 921--950
2016
-
[117]
Xue, F., Zhang, Y., Zhou, W., Fu, H., and Qu, A. (2022). Multicategory angle-based learning for estimating optimal dynamic treatment regimes with censored data. Journal of the American Statistical Association\/ , 117 (539), 1438--1451
2022
-
[118]
Yang, Y. (1999). Minimax nonparametric classification. I. rates of convergence. IEEE Transactions on Information Theory\/ , 45 (7), 2271--2284
1999
-
[119]
Zajonc, T. (2012). Bayesian inference for dynamic treatment regimes: Mobility, equity, and efficiency in student tracking. Journal of the American Statistical Association\/ , 107 (497), 80--92
2012
-
[120]
and Liu, Y
Zhang, C. and Liu, Y. (2014). Multicategory angle-based large-margin classification. Biometrika\/ , 101 (3), 625--640
2014
-
[121]
Zhang, C., Chen, J., Fu, H., He, X., Zhao, Y.-Q., and Liu, Y. (2020). Multicategory outcome weighted margin-based learning for estimating individualized treatment rules. Statistica sinica\/ , 30 , 1857
2020
-
[122]
and Tchetgen, E
Zhang, J. and Tchetgen, E. T. (2024). On identification of dynamic treatment regimes with proxies of hidden confounders. arXiv preprint arXiv:2402.14942\/
2024 arXiv
-
[123]
Zhang, T. (2004). Statistical analysis of some multi-category large margin classification methods. Journal of Machine Learning Research\/ , 5 (Oct), 1225--1251
2004
-
[124]
B., Davidian, M., and Tsiatis, A
Zhang, Y., Laber, E. B., Davidian, M., and Tsiatis, A. A. (2018). Interpretable dynamic treatment regimes. Journal of the American Statistical Association\/ , 113 (524), 1541--1549
2018
-
[125]
Zhang, Z., Jordan, M., Li, W.-J., and Yeung, D.-Y. (2009). Coherence functions for multicategory margin-based classification methods. In Artificial Intelligence and Statistics\/ , pages 647--654. PMLR
2009
-
[126]
J., and Kosorok, M
Zhao, Y.-Q., Zeng, D., Rush, A. J., and Kosorok, M. R. (2012). Estimating individualized treatment rules using outcome weighted learning. Journal of the American Statistical Association\/ , 107 , 1106–1118
2012
-
[127]
B., and Kosorok, M
Zhao, Y.-Q., Zeng, D., Laber, E. B., and Kosorok, M. R. (2015). New statistical learning methods for estimating optimal dynamic treatment regimes. Journal of the American Statistical Association\/ , 110 , 583--598
2015
-
[128]
B., Ning, Y., Saha, S., and Sands, B
Zhao, Y.-Q., Laber, E. B., Ning, Y., Saha, S., and Sands, B. E. (2019). Efficient augmentation and relaxation learning for individualized treatment rules using observational data. Journal of Machine Learning Research\/ , 20 (48), 1--23
2019
-
[129]
Zhao, Y.-Q., Zhu, R., Chen, G., and Zheng, Y. (2020). Constructing dynamic treatment regimes with shared parameters for censored data. Statistics in medicine\/ , 39 (9), 1250--1263
2020
-
[130]
Zhou, Z., Athey, S., and Wager, S. (2022). Offline multi-action policy learning: Generalization and optimization. Operations Research\/
2022
-
[131]
Zou, H., Zhu, J., and Hastie, T. (2008). New multicategory boosting algorithms based on multicategory Fisher -consistent losses. The Annals of Applied Statistics\/ , 2 (4), 1290
2008
-
[132]
, " * write output.state after.block = add.period write newline
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year label extra.label sort.label INTEGERS output.state before.all mid.sentence after.sentence after.block ...
-
[133]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.