Pith. sign in

REVIEW 2 major objections 7 minor 133 references

On Fisher Consistency of Surrogate Losses for Optimal Dynamic Treatment Regimes with Multiple Categorical Treatments per Stage

T0 review · 2 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Within the class of non-negative, stagewise separable surrogate losses, Fisher consistency for optimal dynamic treatment regimes holds exactly when each single-stage component satisfies Condition N1 and every component at stages t ≥ 2…

desk verdict A substantial sufficiency-side contribution, but the paper's headline necessary-and-sufficient claim is false as stated for P0: a constant shift of their own surrogates refutes necessity. read the letter →

arxiv 2505.17285 v1 pith:BLOUEL5Q submitted 2025-05-22 math.ST stat.TH

classification math.STstat.TH MSC 62C0562H3062L20
keywords dynamictreatmentregimesFisherconsistencysurrogatelossesclassificationcalibrationmulticlassnon-convexoptimizationsequentialdecisionmaking
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper attempts to characterize, for the first time in the dynamic treatment regime (DTR) literature, exactly which surrogate losses can be safely substituted for the discontinuous sequential decision loss when treatments are categorical and chosen over multiple stages. It claims that within the natural class of non-negative, stagewise separable surrogates (products of single-stage surrogates), Fisher consistency holds if and only if every single-stage component is Fisher consistent for single-stage decisions (Condition N1) and every component for stages beyond the first has a convex-hull image set large enough to contain a scaled simplex (Condition N2). If true, this pins down when maximizing the surrogate over all score functions recovers an optimal DTR, and it explains why convex surrogates generally fail: the theorem shows that smooth concave permutation-equivariant surrogates cannot be Fisher consistent beyond a single stage. The paper also introduces SDSS, a gradient-based method using smooth non-concave surrogates that satisfy the conditions, with regret bounds matching minimax rates for binary classification under small-noise assumptions.

What carries the argument

The load-bearing object is the single-stage functional Ψ*_t(p) = sup_{x∈$R^{{k_t}}$} Σ_{a∈[k_t]} p_a φ_t(x;a), the support function of the surrogate's image set V_φ = {φ(x;·) : x∈R^k}. Condition N1 states that maximizers of Ψ_t(·;p) always have pred(x) in the argmax of p; Condition N2 states that Ψ*_t(p) = C_φ max(p), which by Lemma 4.2 is equivalent to the scaled simplex C_φ $S^{{k-1}}$ lying inside conv(V_φ). The proof of the main theorems shows that when N1 and N2 hold, the surrogate regret V^ψ_* − V^ψ(f) is bounded below by a linear multiple of the true regret V* − V(f), and conversely constructs distributions showing that when these conditions fail, a maximizing score sequence can approach the surrogate optimum while the induced policies miss the optimal treatment regime.

What would settle it

Construct a product-form surrogate whose first-stage component satisfies Condition N1 and whose later-stage component satisfies N1 but not N2 (for instance, a shifted surrogate from Lemma 4.3), then search for a distribution satisfying Assumptions I–V with Y_t > C_min > 0 under which maximizing the surrogate value still yields an optimal DTR; finding such a pair would refute the necessity direction for P0, and a smooth concave PERM loss that is Fisher consistent for T = 2 would refute Theorem 3.1.

Watch

Extended reading notes

Core claim

The central discovery is an exact necessary-and-sufficient condition for Fisher consistency of separable surrogates in general multi-stage DTR problems. A surrogate of the product form ψ = ∏_t φ_t is Fisher consistent with respect to all distributions satisfying Assumptions I–V if and only if each single-stage φ_t satisfies Condition N1 (the natural single-stage calibration condition: a score vector that does not point to an optimal treatment cannot maximize the surrogate risk) and, for every t ≥ 2, φ_t additionally satisfies Condition N2, meaning its sup-functional Ψ*_t(p) = sup_x Σ_i p_i φ_t(x;i) equals C_φ max(p). Geometrically, N2 forces the closed convex hull of the image set {φ(x; ·)} to contain a scaled probability simplex, the image hull of the original discontinuous loss. The paper proves sufficiency for all T and necessity under non-negative rewards (Assumptions I–IV with Y_t ≥ 0), and argues the same conditions should hold for the strictly positive reward class P0.

Load-bearing premise

The necessity half of the main theorem is proved only for distributions whose rewards are non-negative (Y_t ≥ 0), and the paper asserts the same conditions are necessary for the strictly-positive-reward class P0 without proving that transfer; if a separable surrogate were Fisher consistent on P0 while violating Condition N2, the 'if and only if' claim as stated would fail.

Editorial extensions

If this is right

  • Maximizing any product-form surrogate that satisfies Conditions N1 and N2 recovers an optimal dynamic treatment regime exactly, while violating either condition allows a maximizing sequence to converge to a suboptimal policy.
  • No smooth concave permutation-equivariant and relative-margin-based (PERM) surrogate can be Fisher consistent for two or more stages, so convexification of simultaneous direct search is impossible within that class.
  • The proposed kernel-based and product-based surrogates satisfy N1–N3, yielding smooth non-concave losses with a linear regret constant and a regret decay rate n^{-(1+α)/(2+α+q/β)} up to log factors under small-noise and Hölder-smooth blip assumptions.
  • For restricted policy classes such as linear policies, Conditions N1–N3 still deliver regret bounds when the optimal policy lies inside the class, giving a formal guarantee for interpretable SDSS implementations.
  • Shifted single-stage surrogates that satisfy Condition N1 but fail Condition N2 are provably not Fisher consistent for T ≥ 2, so the multi-stage setting demands surrogates geometrically closer to the original discontinuous loss.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The characterization points to a 'no free lunch' geometry: any Fisher consistent surrogate for multi-stage DTR must have an image hull large enough to contain a scaled simplex, a strictly stronger requirement than single-stage calibration, which explains why the multi-stage problem resists convexification.
  • The necessity-direction proof gap (Y_t ≥ 0 versus Y_t > C_min > 0) might be patchable by a location-shift argument, but the paper does not supply it; a counterexample on the strictly positive class P0 would reveal a genuine boundary phenomenon where the iff statement fails as written.
  • The SDSS regret analysis treats optimization error as an additive black box, yet the paper's own toy example shows gradient methods can get trapped in suboptimal plateaus, so practical performance depends critically on the heuristics (restarts, minibatching, learning-rate scheduling) rather than on the surrogate conditions alone.
  • The product-based surrogate has a slower approximation error (δ^α) than the symmetric kernel-based surrogate (δ^{1+α}), so among Fisher consistent surrogates the choice affects the achievable rate—an implicit trade-off the paper notes but leaves for future selection.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. This paper studies the sequential DTR classification problem with T stages and k_t ≥ 2 treatment options per stage under known propensity scores, and it contributes both theory and methodology. The central theoretical claim is a characterization of Fisher consistency within the class of non-negative stagewise separable surrogates ψ(x₁,…,x_T; a₁,…,a_T) = ∏_{t=1}^T φ_t(x_t; a_t): Theorem 4.1 shows that Condition N1 for every stage plus Condition N2 for every stage t ≥ 2 are sufficient for Fisher consistency on the primary class P₀ (Assumptions I–V with Y_t > C_min > 0), and Theorem 4.2 shows that F1–F2 are necessary for Fisher consistency on the larger class of distributions with Y_t ≥ 0; the paper then refers to F1–F2 as the necessary and sufficient conditions for the separable class. Additional contributions include an impossibility theorem for smooth concave PERM surrogates (Theorem 3.1), explicit smooth non-convex surrogates (kernel- and product-based) with closed-form constants, the SDSS optimization algorithm, a regret bound under a Tsybakov small-noise condition (Theorem 6.2, Corollary 1), and an empirical evaluation on simulations and on MIMIC-III sepsis data.

Significance. The paper's strengths are substantial. The sufficient conditions and surrogate constructions are concrete, closed-form, and checkable (explicit χ_φ, J_t, and C* constants), which makes the N1–N2 conditions readily falsifiable against new surrogate proposals; the PERM impossibility result is non-trivial and goes beyond prior binary-treatment work; and the regret analysis attains a rate matching the binary-classification benchmark under comparable Tsybakov-type conditions. The empirical study is careful, with publicly available code, standard comparators (Q-learning and ACWL), and a sensible AIPW-based evaluation on MIMIC-III EHR data. If the headline characterization were fully established for P₀, this would be the first necessary-and-sufficient DTR Fisher-consistency result in a non-trivial surrogate class. The main weakness is that the necessity half of the characterization is proved only for the larger Y_t ≥ 0 class, not for the paper's primary class P₀, while the Abstract and Introduction present the result as a complete necessary-and-sufficient characterization; the body of the paper acknowledges this gap explicitly but then sets it aside.

major comments (2)
  1. [Section 4, Theorems 4.1–4.2; Abstract] The headline claim of a necessary-and-sufficient characterization for the paper's primary class P₀ is not established by the theorems as stated. Theorem 4.2 proves necessity of F1–F2 under the hypothesis that ψ is Fisher consistent on the strictly larger class of distributions satisfying Assumptions I–IV with Y_t ≥ 0. Because P₀ (Assumption V: Y_t > C_min > 0) is a proper subset of that class, Fisher consistency on P₀ is a weaker property than Fisher consistency on the larger class; the implication 'FC on {Y_t ≥ 0} implies F1+F2' therefore does not transfer downward to P₀. The paper's own text acknowledges this ('It may be possible to show that F1 and F2 are necessary for Fisher consistency even under P₀... considerably more technical') but then states 'Despite this minor discrepancy, we will refer to F1 and F2 as the necessary and sufficient conditions...'. The discrepancy is not minor for the central claim: the necessity of F1+F2 for P₀ remains an open proof obligation, and the Abstract's claim to establish 'necessary and sufficient conditions for DTR Fisher consistency within the class of non-negative, stagewise separable surrogate losses' overstates what Theorems 4.1–4.2 prove. The authors should either supply the P₀-necessity proof or visibly restrict the characterization to the Y_t ≥ 0 class and present the P₀ direction as a conjecture.
  2. [Section 4, Lemma 4.3; Section 2, Definition 2.1] A natural attempt to falsify the P₀-necessity claim—adding a positive constant to a surrogate satisfying N1+N2, as permitted by Lemma 4.3, and noting that the shift violates N2—does not produce a counterexample, because the shifted surrogate is not Fisher consistent on all of P₀. Concretely, take T = 2, k₁ = k₂ = 2, no covariates, π₁ = π₂ = 1/2, deterministic outcomes with E[Y₁+Y₂ | A₁,A₂] = [[10, 9], [1, 10.1]] (realizable with Y₁ ≡ 0.1 and Y₂ ∈ {9.9, 8.9, 0.9, 10.0}, so the distribution lies in P₀), and the logistic shifted surrogate φ(x; i) = σ(x_i − x_{3−i}) + γ with γ = 0.1. Writing s = σ(f₁₁ − f₁₂), u = σ(f₂(1)₁ − f₂(1)₂), v = σ(f₂(2)₁ − f₂(2)₂), direct computation gives V_ψ(s,u,v) = (s + 0.1)(u + 10.9) + (1.1 − s)(11.21 − 9.1v). The supremum, 14.211, is approached only as s → 1, u → 1, v → 0, so every maximizing sequence has pred(f₁) = 1 and pred(f₂(1)) = 1; consequently V(f_m) → W(1,1) = 10, while V* = 10.1 is attained only by d₁ = 2. The shifted surrogate thus fails Fisher consistency at this P ∈ P₀. The proof gap in the necessity direction for P₀ stands, but the available evidence points to the characterization being true rather than false; the manuscript should close the gap with a proof or state the P₀-necessity as an open problem.
minor comments (7)
  1. [Section 3, Example 1] The sentence after Result 1, 'the loss in (3.1) can not be Fisher inconsistent', states the opposite of the intended conclusion and should read 'cannot be Fisher consistent'.
  2. [Section 1.1.1] The heading 'Thoretical results on Fisher consistency' contains a typo; it should read 'Theoretical results on Fisher consistency'.
  3. [Equation (5.4)] Equation (5.4) contains the stray symbol '≡='; a single equals sign is evidently intended.
  4. [Table 1] The table conflates the score-vector columns with the argmax-set column; since the pred link returns max(argmax), the all-zero score vectors in Settings 1–2 yield pred = 3, and this tie-breaking should be stated explicitly where the table is discussed.
  5. [Section 4.3, Lemma 4.8] In the product-based template, the i = 1 case is written as ∏_{j∈[k]} τ(y_j), but (4.11) and the definition Δx₁ = (x₁ − x₂, …, x₁ − x_k) require the product to run over j ∈ [k] \ {1}, excluding the own-margin term; the same indexing issue affects (4.14).
  6. [Section 4.0.2] Immediately after Theorem 4.1, the sentence 'The proof of Theorem 4.2 can be found in Supplement S7' must refer to Theorem 4.1, since the following sentence describes the sufficiency construction.
  7. [Section 2, Assumption V] The 'without loss of generality' location-shift justification is valid for the value function because V'(d) = V(d) + T·c preserves the argmax, but the text should state explicitly that the surrogate analysis is applied to the shifted outcomes and that the shift changes the finite-sample surrogate objective, a point the paper itself later raises in Section 9.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity. The necessary-and-sufficient characterization is derived in the paper from the definitions and from standard external calibration results; self-citations are contextual rather than load-bearing. An acknowledged proof gap in the necessity direction (Y_t >= 0 vs P0) is a correctness concern, not a circularity.

full rationale

The paper's central claim is a characterization theorem: within the class of non-negative stagewise separable surrogates, DTR Fisher consistency holds if and only if each single-stage component satisfies Condition N1 and, for t >= 2, Condition N2. The proof derives these conditions from the definition of Fisher consistency via constructed 'good sets' of distributions, rather than assuming the conditions or fitting them to the target. No numeric parameters are fitted and then renamed as predictions; the kernel-based and product-based surrogates are defined in closed form and then verified to satisfy the stated conditions. The necessity proof in Theorem 4.2 is conducted for the broader class of distributions with Y_t >= 0, and the paper explicitly acknowledges that transferring necessity to P0 (Y_t > C_min > 0) is not proved: 'It may be possible to show that F1 and F2 are necessary for Fisher consistency even under P0... the corresponding proof would be considerably more technical.' This is a proof gap and a possible correctness defect, but it is not circularity: it does not reduce the conclusion to an input by construction. Self-citations to Laha et al. (2024a) are used for comparison, for examples, and for prior binary-treatment results; the main separable-surrogate characterization is proved in this paper from the definitions and from standard external multiclass calibration results (Zhang 2004; Tewari and Bartlett 2007; Wang and Scott 2023a,b), which are independent support. The Theorem 4.1 sufficient direction and the regret bounds are likewise derived rather than assumed. No equation in the paper is equivalent to its own input by construction, so the appropriate circularity score is 1.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The paper introduces no new physical entities; the kernel-based and product-based surrogates are mathematical constructions rather than invented entities. The central characterization rests on standard DTR identifiability assumptions, a small-noise condition, smoothness conditions for rates, and existing multiclass calibration theorems. No numeric parameters are fitted to data in deriving the theoretical results; the constant C in the surrogates is a scale factor that cancels in regret bounds.

assumptions (6)
  • domain assumption Identifiability assumptions I-III: positivity, consistency, and sequential ignorability for the IPW representation of the value function (Equation 2.2).
    Required to write V(d) as an expectation under the observed distribution. Standard in DTR literature, stated in Section 2.
  • domain assumption Assumptions IV-V: outcomes are bounded above and can be shifted so that Y_t > C_min > 0 for all t.
    Used throughout the Fisher consistency proofs and to ensure surrogates behave well; location shift is described as a WLOG step in Section 2.
  • domain assumption Tsybakov small noise condition (Assumption 1) on the gap in optimal Q-functions.
    Needed for the sharp regret decay rates in Section 6; standard in classification and DTR regret analysis.
  • domain assumption Hölder smoothness of blip functions (Assumption 2) for the neural network rate in Corollary 1.
    Needed to bound the approximation error of neural network policy classes; stated in Section 6.3.1.
  • standard math Known multiclass Fisher consistency characterizations (Tewari and Bartlett 2007, Zhang 2004, Wang and Scott 2023a,b).
    Used to define Condition N1 as the single-stage condition and to identify single-stage Fisher consistent losses such as the exponential loss.
  • standard math Convex analysis facts: support functions, closed convex hulls, proper and closed convex functions, and strict convexity conditions on the negative template.
    Used extensively in the proofs of Theorem 3.1 and the Lemma 4.2 geometric characterization of Condition N2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On Fisher Consistency of Surrogate Losses for Optimal Dynamic Treatment Regimes with Multiple Categorical Treatments per Stage." pith.science (2026). https://pith.science/paper/BLOUEL5Q

@misc{pith2026250517285,
  author       = {Pith},
  title        = {Pith review of: On Fisher Consistency of Surrogate Losses for Optimal Dynamic Treatment Regimes with Multiple Categorical Treatments per Stage},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BLOUEL5Q}},
  note         = {Machine review of arXiv:2505.17285}
}
read the original abstract

Patients with chronic diseases often receive treatments at multiple time points, or stages. Our goal is to learn the optimal dynamic treatment regime (DTR) from longitudinal patient data. When both the number of stages and the number of treatment levels per stage are arbitrary, estimating the optimal DTR reduces to a sequential, weighted, multiclass classification problem (Kosorok and Laber, 2019). In this paper, we aim to solve this classification problem simultaneously across all stages using Fisher consistent surrogate losses. Although computationally feasible Fisher consistent surrogates exist in special cases, e.g., the binary treatment setting, a unified theory of Fisher consistency remains largely unexplored. We establish necessary and sufficient conditions for DTR Fisher consistency within the class of non-negative, stagewise separable surrogate losses. To our knowledge, this is the first result in the DTR literature to provide necessary conditions for Fisher consistency within a non-trivial surrogate class. Furthermore, we show that many convex surrogate losses fail to be Fisher consistent for the DTR classification problem, and we formally establish this inconsistency for smooth, permutation equivariant, and relative-margin-based convex losses. Building on this, we propose SDSS (Simultaneous Direct Search with Surrogates), which uses smooth, non-concave surrogate losses to learn the optimal DTR. We develop a computationally efficient, gradient-based algorithm for SDSS. When the optimization error is small, we establish a sharp upper bound on SDSS's regret decay rate. We evaluate the numerical performance of SDSS through simulations and demonstrate its real-world applicability by estimating optimal fluid resuscitation strategies for severe septic patients using electronic health record data.

Figures

Figures reproduced from arXiv: 2505.17285 by the authors.

Figure 1
Figure 1. Plot of Γ(x, y; 1) when k = 3. For the product-based surrogate, Γ is as in (4.13), and τ (x) = (1 + tanh(x))/2, which is the distribution function of the centered logistic distribution with scale 2. For kernel-based surrogates, the template Γ is provided in (4.14). Its closed formulas for the logistic and Gumbel densities are provided in Supplement S8.5.1. Relative margin-based surrogate losses are quite common in m… view at source ↗
Figure 2
Figure 2. Plots related to Vb ψ,rel for the toy example in Section 5.1. The formula of Vb ψ,rel(x, y) is provided in (5.4) and the toy data is provided in [PITH_FULL_IMAGE:figures/full_fig_p026_2.png] view at source ↗
Figure 3
Figure 3. Gradient descent for the toy data in Section 5.1 The plots display the paths traced by iterates initiated from 6 different initialization points for (a) vanilla gradient descent, (b) ADAM (without minibatching), (c) SGD, and (d) ADAM with SGD. The white circle and the solid black rectangle mark the starting point and the end point of the paths, respectively. Although the algorithms minimize −Vb ψ,rel, the paths are … view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Optimal treatment assignments for Scheme 2 as a function of OT 1 1p. As mentioned in Section 7.1, the optimal treatment assignments in Scheme 2 are functions of OT 1 1p only. The X-axis represents OT 1 1p, while the Y-axis indicates the corresponding optimal treatment.…
Figure 5
Figure 5. Figure 5: Boxplot of value functions: Boxplot of estimated value function of different methods for 100 replications. The value functions in Scheme 1 are scaled by 10−2 for better visual comparability. The X-axis indicates the sample size and the Y-axis represents the average val…
Figure 6
Figure 6. Figure 6: provides the boxplots of the value function estimates obtained from the 100 replications. As in Section 7.1, ACWL was excluded from linear policy computations. The Python code for the data application is available in the GitHub repository cited as Chapagain (2024a). (a…
Figure 7
Figure 7. Figure 7: Plot of mean treatment level vs SOFA scores in our data application for non-linear policies. We partition the SOFA score into four clinically meaningful bins, 0–4, 5–8, 9–12, and ≥13, corresponding to mild, moderate, severe, and critical organ dysfunction. Treatment le…
Figure 8
Figure 8. Figure 8: Value function of SDSS with and without ensembling under Scheme 2. Ensembling is performed as described in Section 7.1. The regular SDSS (colored blue) uses τ (x) = 1 + x/√ 1 + x 2. The left panel shows results for linear policies; the right panel shows results for non…
Figure 9
Figure 9. Figure 9: Surface plot of ψ defined in (S4.8) for the last two settings in [PITH_FULL_IMAGE:figures/full_fig_p048_9.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

133 extracted references · 76 canonical work pages

  1. [1]

    Akhtar, M., Tanveer, M., and Arshad, M. (2024). Hawkeye: advancing robust regression with bounded, smooth, and insensitive loss function. arXiv preprint arXiv:2401.16785\/

  2. [2]

    and Wager, S

    Athey, S. and Wager, S. (2021). Policy learning with observational data. Econometrica\/ , 89 (1), 133--161

  3. [3]

    and Tsybakov, A

    Audibert, J.-Y. and Tsybakov, A. B. (2007). Fast learning rates for plug-in classifiers. The Annals of statistics\/ , 35 (2), 608--633

  4. [4]

    L., Jordan, M

    Bartlett, P. L., Jordan, M. I., and McAuliffe, J. D. (2006). Convexity, classification, and risk bounds. Journal of the American Statistical Association\/ , 101 (473), 138--156

  5. [5]

    Beijbom, O., Saberian, M., Kriegman, D., and Vasconcelos, N. (2014). Guess-averse loss functions for cost-sensitive multiclass boosting. In International Conference on Machine Learning\/ , pages 586--594. PMLR

  6. [6]

    and Kallus, N

    Bennett, A. and Kallus, N. (2020). Efficient policy learning from surrogate-loss classification reductions. In International Conference on Machine Learning\/ , pages 788--798. PMLR

  7. [7]

    Bhattacharyya, A., Ghoshal, S., and Saket, R. (2018). Hardness of learning noisy halfspaces using polynomial thresholds. In Conference On Learning Theory\/ , pages 876--917. PMLR

  8. [8]

    Billingsley, P. (2017). Probability and measure\/ . John Wiley & Sons

Show all 133 references
  1. [9]

    E., and Nocedal, J

    Bottou, L., Curtis, F. E., and Nocedal, J. (2018). Optimization methods for large-scale machine learning. SIAM review\/ , 60 (2), 223--311

  2. [10]

    and Vandenberghe, L

    Boyd, S. and Vandenberghe, L. (2004). Convex Optimization\/ . Cambridge University Press, New York, NY, USA

  3. [11]

    Calauzenes, C., Usunier, N., and Gallinari, P. (2012). On the (non-) existence of convex, calibrated surrogate losses for ranking. Advances in Neural Information Processing Systems 25 (NIPS 2012)\/ , pages 197--205

  4. [12]

    and Moodie, E

    Chakraborty, B. and Moodie, E. E. (2013). Statistical methods for dynamic treatment regimes\/ , volume 2. Springer

  5. [13]

    Chapagain, N. (2024a). Data application code for SDSS method. https://github.com/nilson01/ApplicationFin . GitHub repository

  6. [14]

    Chapagain, N. (2024b). Simulation code for SDSS method. https://github.com/nilson01/SimulationDirectSearch. GitHub repository

  7. [15]

    Charoenphakdee, N., Lee, J., and Sugiyama, M. (2019). On symmetric losses for learning from corrupted labels. In International Conference on Machine Learning\/ , pages 961--970. PMLR

  8. [16]

    R., and Liu, Y

    Chen, J., Fu, H., He, X., Kosorok, M. R., and Liu, Y. (2018). Estimating individualized treatment rules for ordinal treatments. Biometrics\/ , 74 (3), 924--933

  9. [17]

    Chen, Y., Chi, Y., Fan, J., and Ma, C. (2019). Gradient descent with random initialization: Fast global convergence for nonconvex phase retrieval. Mathematical Programming\/ , 176 , 5--37

  10. [18]

    N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y

    Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014). Identifying and attacking the saddle point problem in high-dimensional non-convex optimization. In Advances in Neural Information Processing Systems\/ , pages 2933--2941

  11. [19]

    J., and Chakraborty, B

    Deliu, N., Williams, J. J., and Chakraborty, B. (2022). Reinforcement learning in modern biostatistics: constructing optimal adaptive interventions. arXiv preprint arXiv:2203.02605\/

  12. [20]

    N., Dao, C

    Do, S. N., Dao, C. X., Nguyen, T. A., Nguyen, M. H., Pham, D. T., Nguyen, N. T., Huynh, D. Q., Hoang, Q. T. A., Van Bui, C., Vu, T. D., et al. (2023). Sequential organ failure assessment (sofa) score for predicting mortality in patients with sepsis in vietnamese intensive care...

  13. [21]

    Doss, C. R. and Wellner, J. A. (2016). Global rates of convergence of the MLEs of log-concave and s -concave densities. Ann. Statist. , 44 , 954--981

  14. [22]

    Doss, C. R. and Wellner, J. A. (2019). Univariate log-concave density estimation with symmetry or modal constraints. Electron. J. Stat. , 13 , 2391--2461

  15. [23]

    C., Mackey, L

    Duchi, J. C., Mackey, L. W., and Jordan, M. I. (2010). On the consistency of ranking algorithms. In ICML\/

  16. [24]

    E., and Telgarsky, M

    Dud \' k, M., Schapire, R. E., and Telgarsky, M. (2022). Convex analysis at infinity: An introduction to astral space. arXiv preprint arXiv:2205.03260\/

  17. [25]

    Fathony, R., Liu, A., Asif, K., and Ziebart, B. (2016). Adversarial multiclass classification: A risk minimization perspective. Advances in Neural Information Processing Systems\/ , 29

  18. [26]

    Feng, H., Ning, Y., and Zhao, J. (2022). Nonregular and minimax estimation of individualized thresholds in high dimension with binary responses. The Annals of Statistics\/ , 50 (4), 2284--2305

  19. [27]

    Finocchiaro, J., Frongillo, R., and Waggoner, B. (2019). An embedding framework for consistent polyhedral surrogates. Advances in neural information processing systems\/ , 32

  20. [28]

    and Zhou, Z.-H

    Gao, W. and Zhou, Z.-H. (2011). On the consistency of multi-label learning. In Proceedings of the 24th annual conference on learning theory\/ , pages 341--358

  21. [29]

    and Nickl, R

    Giné, E. and Nickl, R. (2015). Mathematical Foundations of Infinite-Dimensional Statistical Models\/ , volume 40 of Cambridge series in statistical and probabilistic mathematics\/ . Cambridge University Press, Cambridge

  22. [30]

    Glasmachers, T., Igel, C., et al. (2016). A unified view on multi-class support vector classification. Journal of Machine Learning Research\/ , 17 (45), 1--32

  23. [31]

    and Bengio, Y

    Glorot, X. and Bengio, Y. (2010). Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics\/ , pages 249--256. JMLR Workshop and Conference Proceedings

  24. [32]

    Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep Learning\/ . MIT Press. http://www.deeplearningbook.org

  25. [33]

    A., and Davidian, M

    Hager, R., Tsiatis, A. A., and Davidian, M. (2018). Optimal two-stage dynamic treatment regimes from a classification perspective with censored survival data. Biometrics\/ , 74 (4), 1180--1192

  26. [34]

    Han, S. (2021). Identification in nonparametric models for dynamic treatment effects. Journal of Econometrics\/ , 225 (2), 132--147

  27. [35]

    He, K., Zhang, X., Ren, S., and Sun, J. (2015). Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE International Conference on Computer Vision\/ , pages 1026--1034

  28. [36]

    W., and Ridder, G

    Hirano, K., Imbens, G. W., and Ridder, G. (2003). Efficient estimation of average treatment effects using the estimated propensity score. Econometrica\/ , 71 (4), 1161--1189

  29. [37]

    and Lemar \'e chal, C

    Hiriart-Urruty, J.-B. and Lemar \'e chal, C. (2004). Fundamentals of convex analysis\/ . Springer Science & Business Media

  30. [38]

    Hochreiter, S., Bengio, Y., Frasconi, P., Schmidhuber, J., et al. (2001). Gradient flow in recurrent nets: the difficulty of learning long-term dependencies

  31. [39]

    Horn, R. A. and Johnson, C. R. (2012). Matrix analysis\/ . Cambridge university press

  32. [40]

    Horowitz, J. L. (1992). A smoothed maximum score estimator for the binary response model. Econometrica: journal of the Econometric Society\/ , pages 505--531

  33. [41]

    Jiang, B., Song, R., Li, J., and Zeng, D. (2019). Entropy learning for dynamic treatment regimes. Statistica Sinica\/ , 29 (4), 1633

  34. [42]

    and Li, L

    Jiang, N. and Li, L. (2016). Doubly robust off-policy value evaluation for reinforcement learning. In International Conference on Machine Learning\/ , pages 652--661. PMLR

  35. [43]

    Jin, C., Liu, Q., and Miryoosefi, S. (2021). Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms. Advances in neural information processing systems\/ , 34 , 13406--13418

  36. [44]

    and Pollard, T

    Johnson, A. and Pollard, T. (2018). sepsis3-mimic

  37. [45]

    Johnson, A., Pollard, T., Shen, L., Lehman, L., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Celi, L., and Mark, R. (2016). Mimic-iii, a freely accessible critical care database. Scientific Data\/ , 3 , 160035

  38. [46]

    E., Aboab, J., Raffa, J

    Johnson, A. E., Aboab, J., Raffa, J. D., Pollard, T. J., Deliberato, R. O., Celi, L. A., and Stone, D. J. (2018). A comparative analysis of sepsis identification methods in an electronic database. Critical care medicine\/ , 46 (4), 494

  39. [47]

    and Uehara, M

    Kallus, N. and Uehara, M. (2020). Double reinforcement learning for efficient off-policy evaluation in markov decision processes. Journal of Machine Learning Research\/ , 21 (167)

  40. [48]

    Kingma, D. P. and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980\/

  41. [49]

    Koltchinskii, V. (2009). 2008 saint flour lectures oracle inequalities in empirical risk minimization and sparse recovery problems

  42. [50]

    Komorowski, M., Celi, L., Badawi, O., Gordon, A., and Faisal, A. (2018). The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care. Nature Medicine\/ , 24 (11), 1716--1720

  43. [51]

    Kosorok, M. R. and Laber, E. B. (2019). Precision medicine. Annual Review of Statistics and its Application \/ , 6 , 263--286

  44. [52]

    B., Lizotte, D

    Laber, E. B., Lizotte, D. J., Qian, M., Pelham, W. E., and Murphy, S. A. (2014). Dynamic treatment regimes: Technical challenges and applications. Electronic journal of statistics\/ , 8 (1), 1225

  45. [53]

    Laha, N. (2021). Adaptive estimation in symmetric location model under log-concavity constraint. Electronic Journal of Statistics \/ , 15 , 2939--3014

  46. [54]

    Laha, N., Sonabend-W., A., Mukherjee, R., and Cai, T. (2024a). Finding the optimal dynamic treatment regimes using smooth fisher consistent surrogate loss. Annals of Statistics\/ , 52 (2), 679--707

  47. [55]

    Laha, N., Sonabend-W, A., Mukherjee, R., and Cai, T. (2024b). Finding the optimal dynamic treatment regimes using smooth fisher consistent surrogate loss

  48. [56]

    M., and De Backer, D

    Lat, I., Coopersmith, C. M., and De Backer, D. (2021). The surviving sepsis campaign: fluid resuscitation and vasopressor therapy research priorities in adult patients. Intensive Care Medicine Experimental\/ , 9 (1), 1--16

  49. [57]

    Lee, Y., Lin, Y., and Wahba, G. (2004). Multicategory support vector machines: Theory and application to the classification of microarray data and satellite radiance data. Journal of the American Statistical Association\/ , 99 (465), 67--81

  50. [58]

    Liu, M., Shen, X., and Pan, W. (2021). Outcome weighted -learning for individualized treatment rules. Stat\/ , 10 (1), e343

  51. [59]

    Liu, M., Wang, Y., Fu, H., and Zeng, D. (2024a). Controlling cumulative adverse risk in learning optimal dynamic treatment regimens. Journal of the American Statistical Association\/ , 119 (548), 2622--2633

  52. [60]

    Liu, M., Wang, Y., Fu, H., and Zeng, D. (2024b). Learning optimal dynamic treatment regimens subject to stagewise risk controls. Journal of Machine Learning Research\/ , 25 (128), 1--64

  53. [61]

    Liu, Y. (2007). Fisher consistency of multicategory support vector machines. In Artificial intelligence and statistics\/ , pages 291--298

  54. [62]

    and Shen, X

    Liu, Y. and Shen, X. (2006). Multicategory -learning. Journal of the American Statistical Association\/ , 101 (474), 500--509

  55. [63]

    R., Zhao, Y., and Zeng, D

    Liu, Y., Wang, Y., Kosorok, M. R., Zhao, Y., and Zeng, D. (2018). Augmented outcome-weighted learning for estimating optimal dynamic treatment regimens. Statistics in medicine\/ , 37 (26), 3776--3788

  56. [64]

    Luedtke, A. R. and Van Der Laan, M. J. (2016). Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy. Annals of statistics\/ , 44 (2), 713

  57. [65]

    Marik, P. (2015). The demise of early goal-directed therapy for severe sepsis and septic shock. Acta Anaesthesiologica Scandinavica\/ , 59 (5), 561--567

  58. [66]

    and Vasconcelos, N

    Masnadi-Shirazi, H. and Vasconcelos, N. (2008). On the design of loss functions for classification: theory, robustness to outliers, and savageboost. Advances in neural information processing systems\/ , 21

  59. [67]

    Meng, H., Zhao, Y.-Q., Fu, H., and Qiao, X. (2020). Near-optimal individualized treatment recommendations. J. Mach. Learn. Res. , 21 , 183--1

  60. [68]

    E., Richardson, T

    Moodie, E. E., Richardson, T. S., and Stephens, D. A. (2007). Demystifying optimal dynamic treatment regimes. Biometrics\/ , 63 (2), 447--455

  61. [69]

    E., Dean, N., and Sun, Y

    Moodie, E. E., Dean, N., and Sun, Y. R. (2014). Q-learning: Flexible learning about useful utilities. Statistics in Biosciences\/ , 6 , 223--243

  62. [70]

    Murphy, S. A. (2003). Optimal dynamic treatment regimes. Journal of the Royal Statistical Society: Series B (Statistical Methodology)\/ , 65 (2), 331--355

  63. [71]

    Murphy, S. A. (2005). A generalization error for Q-learning . Journal Of Machine Learning Research\/ , 6 , 1073--1097

  64. [72]

    A., van der Laan, M

    Murphy, S. A., van der Laan, M. J., and Robins, J. M. (2001a). Marginal mean models for dynamic regimes. Journal of the American Statistical Association\/ , 96 (456), 1410--1423

  65. [73]

    A., van der Laan, M

    Murphy, S. A., van der Laan, M. J., and Robins, J. M. (2001b). Marginal mean models for dynamic regimes. Journal of the American Statistical Association\/ , 96 (456), 1410--1423

  66. [74]

    S., Ravikumar, P

    Natarajan, N., Dhillon, I. S., Ravikumar, P. K., and Tewari, A. (2013). Learning with noisy labels. Advances in neural information processing systems\/ , 26

  67. [75]

    S., and Cai, T

    Neykov, M., Liu, J. S., and Cai, T. (2016). On the characterization of a class of Fisher-consistent loss functions and its application to boosting. The Journal of Machine Learning Research\/ , 17 (1), 2498--2529

  68. [76]

    and Sanner, S

    Nguyen, T. and Sanner, S. (2013). Algorithms for direct 0--1 loss optimization in binary classification. In International conference on machine learning\/ , pages 1085--1093. PMLR

  69. [77]

    Orellana, L., Rotnitzky, A., and Robins, J. M. (2010). Dynamic regime marginal structural mean models for estimation of optimal dynamic treatment regimes, part i: main content. The international journal of biostatistics\/ , 6 (2)

  70. [78]

    Osokin, A., Bach, F., and Lacoste-Julien, S. (2017). On structured prediction theory with calibrated convex surrogate losses. Advances in Neural Information Processing Systems\/ , 30

  71. [79]

    and Zhao, Y.-Q

    Pan, Y. and Zhao, Y.-Q. (2021). Improved doubly robust estimation in learning optimal individualized treatment rules. Journal of the American Statistical Association\/ , 116 (533), 283--294

  72. [80]

    Patrini, G., Rozza, A., Krishna Menon, A., Nock, R., and Qu, L. (2017). Making deep neural networks robust to label noise: A loss correction approach. In Proceedings of the IEEE conference on computer vision and pattern recognition\/ , pages 1944--1952

  73. [81]

    and Murphy, S

    Qian, M. and Murphy, S. A. (2011). Performance guarantees for individualized treatment rules. Annals of statistics\/ , 39 (2), 1180

  74. [82]

    A., Szolovits, P., and Ghassemi, M

    Raghu, A., Komorowski, M., Celi, L. A., Szolovits, P., and Ghassemi, M. (2017). Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach. abs/1705.08422

  75. [83]

    Ramaswamy, H. G. and Agarwal, S. (2016). Convex calibration dimension for multiclass loss matrices. The Journal of Machine Learning Research\/ , 17 (1), 397--441

  76. [84]

    and Zeevi, A

    Rigollet, P. and Zeevi, A. J. (2010). Nonparametric bandits with covariates. In Annual Conference Computational Learning Theory\/

  77. [85]

    Robins, J. M. (1997). Causal inference from complex longitudinal data. In Latent variable modeling and applications to causality\/ , pages 69--117. Springer

  78. [86]

    Robins, J. M. (2004). Optimal structural nested models for optimal sequential decisions. Proceedings of the second Seattle symposium on biostatistics In D. Y. Lin and P. Heagerty (Eds.) (pp. 189--326). New York: Springer

  79. [87]

    M., Rotnitzky, A., and Zhao, L

    Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association\/ , 89 (427), 846--866

  80. [88]

    Rockafellar, R. T. (1970). Convex Analysis \/ . Princeton University Press

  81. [89]

    Rosset, S., Zhu, J., and Hastie, T. (2003). Margin maximizing loss functions. Advances in neural information processing systems\/ , 16

  82. [90]

    and Wellner, J

    Saumard, A. and Wellner, J. A. (2014). Log-concavity and strong log-concavity: a review. Statistics surveys\/ , 8 , 45

  83. [91]

    Nonparametric regression using deep neural networks with relu activation function

    Schmidt-Hieber (2020). Nonparametric regression using deep neural networks with relu activation function. Annals of Statistics\/ , 48 (4), 1875--1897

  84. [92]

    J., Tsiatis, A

    Schulte, P. J., Tsiatis, A. A., Laber, E. B., and Davidian, M. (2014). Q - and A -learning methods for estimating optimal dynamic treatment regimes. Statist. Sci. , 29 (4), 640--661

  85. [93]

    and Wellner, J

    Seregin, A. and Wellner, J. A. (2010). Nonparametric estimation of multivariate convex-transformed densities. Annals of statistics\/ , 38 (6), 3751

  86. [94]

    N., Cai, T., and Mukherjee, R

    Sonabend-W, A., Laha, N., Ananthakrishnan, A. N., Cai, T., and Mukherjee, R. (2023). Semi-supervised off-policy reinforcement learning and value estimation for dynamic treatment regimes. Journal of Machine Learning Research\/ , 24 (323), 1--86

  87. [95]

    Steinwart, I., Scovel, C., et al. (2007). Fast rates for support vector machines using gaussian kernels. The Annals of Statistics\/ , 35 (2), 575--607

  88. [96]

    and Wang, L

    Sun, Y. and Wang, L. (2021). Stochastic tree search for estimating optimal dynamic treatment regimes. Journal of the American Statistical Association\/ , 116 (533), 421--432

  89. [97]

    Sutton, R. S. (2018). Reinforcement learning : an introduction\/ . Adaptive computation and machine learning. The MIT Press, Cambridge, Massachusetts ; London, England, second edition. edition

  90. [98]

    and Wang, L

    Tao, Y. and Wang, L. (2017). Adaptive contrast weighted learning for multi-stage multi-treatment decision-making. Biometrics\/ , 73 (1), 145--155

  91. [99]

    Tao, Y., Wang, L., and Almirall, D. (2018). Tree-based reinforcement learning for estimating optimal dynamic treatment regimes. The annals of applied statistics\/ , 12 (3), 1914

  92. [100]

    and Bartlett, P

    Tewari, A. and Bartlett, P. L. (2007). On the consistency of multiclass classification methods. Journal of Machine Learning Research\/ , 8 (May), 1007--1025

  93. [101]

    and Brunskill, E

    Thomas, P. and Brunskill, E. (2016). Data-efficient off-policy policy evaluation for reinforcement learning. In International Conference on Machine Learning\/ , pages 2139--2148. PMLR

  94. [102]

    Tsiatis, A. (2006). Semiparametric theory and missing data. 73

  95. [103]

    A., Davidian, M., Holloway, S

    Tsiatis, A. A., Davidian, M., Holloway, S. T., and Laber, E. B. (2019). Dynamic treatment regimes: Statistical methods for precision medicine\/ . Chapman and Hall/CRC

  96. [104]

    Tsybakov, A. B. (2004). Optimal aggregation of classifiers in statistical learning. The Annals of Statistics\/ , 32 (1), 135--166

  97. [105]

    and Lederer, J

    Ven, L. and Lederer, J. (2021). Regularization and reparameterization avoid vanishing gradients in sigmoid-type networks. arXiv preprint arXiv:2106.02260\/

  98. [106]

    Wallace, M. P. and Moodie, E. E. (2015). Doubly-robust dynamic treatment regimen estimation via weighted least squares. Biometrics\/ , 71 (3), 636--644

  99. [107]

    Wang, K., Pei, H., Cao, J., and Zhong, P. (2020). Robust regularized extreme learning machine for regression with non-convex loss function via dc program. Journal of the Franklin Institute\/ , 357 (11), 7069--7091

  100. [108]

    and Scott, C

    Wang, Y. and Scott, C. (2023a). On classification-calibration of gamma-phi losses. In The Thirty Sixth Annual Conference on Learning Theory\/ , pages 4929--4951. PMLR

  101. [109]

    and Scott, C

    Wang, Y. and Scott, C. (2023b). Unified binary and multiclass margin-based classification. arXiv preprint arXiv:2311.17778\/

  102. [110]

    Wang, Y., Fu, H., and Zeng, D. (2018). Learning optimal personalized treatment rules in consideration of benefit and risk: with an application to treating type 2 diabetes patients with insulin therapies. Journal of the American Statistical Association\/ , 113 (521), 1--13

  103. [111]

    Watkins, C. J. C. H. (1989). Learning from Delayed Rewards\/ . Ph.D. thesis, King's College, Cambridge, UK

  104. [112]

    Wellner, J. A. (2005). Empirical processes: Theory and applications. Lecture notes

  105. [113]

    Weston, J., Watkins, C., et al. (1999). Support vector machines for multi-class pattern recognition. In Esann\/ , volume 99, pages 219--224

  106. [114]

    and Liu, Y

    Wu, Y. and Liu, Y. (2007). Robust truncated hinge loss support vector machines. Journal of the American Statistical Association\/ , 102 (479), 974--983

  107. [115]

    Xu, T., Wang, J., and Fang, Y. (2014). A model-free estimation for the covariate-adjusted youden index and its associated cut-point. Statistics in medicine\/ , 33 (28), 4963--4974

  108. [116]

    S., and Thall, P

    Xu, Y., M \"u ller, P., Wahed, A. S., and Thall, P. F. (2016). Bayesian nonparametric estimation for dynamic treatment regimes with sequential transition times. Journal of the American Statistical Association\/ , 111 (515), 921--950

  109. [117]

    Xue, F., Zhang, Y., Zhou, W., Fu, H., and Qu, A. (2022). Multicategory angle-based learning for estimating optimal dynamic treatment regimes with censored data. Journal of the American Statistical Association\/ , 117 (539), 1438--1451

  110. [118]

    Yang, Y. (1999). Minimax nonparametric classification. I. rates of convergence. IEEE Transactions on Information Theory\/ , 45 (7), 2271--2284

  111. [119]

    Zajonc, T. (2012). Bayesian inference for dynamic treatment regimes: Mobility, equity, and efficiency in student tracking. Journal of the American Statistical Association\/ , 107 (497), 80--92

  112. [120]

    and Liu, Y

    Zhang, C. and Liu, Y. (2014). Multicategory angle-based large-margin classification. Biometrika\/ , 101 (3), 625--640

  113. [121]

    Zhang, C., Chen, J., Fu, H., He, X., Zhao, Y.-Q., and Liu, Y. (2020). Multicategory outcome weighted margin-based learning for estimating individualized treatment rules. Statistica sinica\/ , 30 , 1857

  114. [122]

    and Tchetgen, E

    Zhang, J. and Tchetgen, E. T. (2024). On identification of dynamic treatment regimes with proxies of hidden confounders. arXiv preprint arXiv:2402.14942\/

  115. [123]

    Zhang, T. (2004). Statistical analysis of some multi-category large margin classification methods. Journal of Machine Learning Research\/ , 5 (Oct), 1225--1251

  116. [124]

    B., Davidian, M., and Tsiatis, A

    Zhang, Y., Laber, E. B., Davidian, M., and Tsiatis, A. A. (2018). Interpretable dynamic treatment regimes. Journal of the American Statistical Association\/ , 113 (524), 1541--1549

  117. [125]

    Zhang, Z., Jordan, M., Li, W.-J., and Yeung, D.-Y. (2009). Coherence functions for multicategory margin-based classification methods. In Artificial Intelligence and Statistics\/ , pages 647--654. PMLR

  118. [126]

    J., and Kosorok, M

    Zhao, Y.-Q., Zeng, D., Rush, A. J., and Kosorok, M. R. (2012). Estimating individualized treatment rules using outcome weighted learning. Journal of the American Statistical Association\/ , 107 , 1106–1118

  119. [127]

    B., and Kosorok, M

    Zhao, Y.-Q., Zeng, D., Laber, E. B., and Kosorok, M. R. (2015). New statistical learning methods for estimating optimal dynamic treatment regimes. Journal of the American Statistical Association\/ , 110 , 583--598

  120. [128]

    B., Ning, Y., Saha, S., and Sands, B

    Zhao, Y.-Q., Laber, E. B., Ning, Y., Saha, S., and Sands, B. E. (2019). Efficient augmentation and relaxation learning for individualized treatment rules using observational data. Journal of Machine Learning Research\/ , 20 (48), 1--23

  121. [129]

    Zhao, Y.-Q., Zhu, R., Chen, G., and Zheng, Y. (2020). Constructing dynamic treatment regimes with shared parameters for censored data. Statistics in medicine\/ , 39 (9), 1250--1263

  122. [130]

    Zhou, Z., Athey, S., and Wager, S. (2022). Offline multi-action policy learning: Generalization and optimization. Operations Research\/

  123. [131]

    Zou, H., Zhu, J., and Hastie, T. (2008). New multicategory boosting algorithms based on multicategory Fisher -consistent losses. The Annals of Applied Statistics\/ , 2 (4), 1290

  124. [132]

    , " * write output.state after.block = add.period write newline

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year label extra.label sort.label INTEGERS output.state before.all mid.sentence after.sentence after.block ...

  125. [133]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.