Pith. sign in

REVIEW 2 major objections 5 minor 2 cited by

Reinforcement Learning in Healthcare: A Survey

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Reinforcement learning works across healthcare, but its clinical value hinges on how states, rewards, and policies are formulated.

desk verdict Broad, well-organized RL-in-healthcare survey whose usefulness does not depend on the unsupported 'first comprehensive survey' claim. read the letter →

arxiv 1908.08796 v4 pith:SOG4EKXG submitted 2019-08-22 cs.LG cs.AI

classification cs.LGcs.AI
keywords reinforcementlearninghealthcaredynamictreatmentregimescriticalcaremedicaldiagnosissurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey paper tries to establish that reinforcement learning (RL) is a broadly applicable and increasingly successful approach for sequential decision-making in healthcare, and that the field's central obstacles are common across applications rather than domain-specific. The authors claim that RL's ability to learn from evaluative, delayed feedback makes it suited to dynamic treatment regimes, automated diagnosis, and resource-scheduling problems. They argue that the main open issues—state and action engineering, reward formulation, policy evaluation, model learning, exploration, and credit assignment—are what currently limit clinical impact, and they call for future work on interpretability, prior knowledge, small data, ambient intelligence, and in-vivo validation.

What carries the argument

The central framework is the Markov decision process (MDP), defined as a 5-tuple $(S, A, P, R, \gamma)$, where an agent chooses actions to maximize discounted cumulative reward. The survey organizes RL techniques along two complementary directions: efficient techniques (experience replay and batch RL, model-based RL, transfer RL) and representational techniques (function approximation, deep RL, multi-objective RL, preference-based RL, inverse RL, factored MDPs, hierarchical RL, POMDPs).

What would settle it

A reader could falsify the survey's breadth claim by identifying a substantial published corpus of RL applications in healthcare that this survey omits, or by conducting a systematic literature search that yields a different distribution of application domains and open problems.

Watch

Extended reading notes

Core claim

The paper's central claim is that RL has been successfully applied across a broad range of healthcare domains, including dynamic treatment regimes for chronic diseases (cancer, diabetes, anemia, HIV, mental illness) and critical care (sepsis, anesthesia, mechanical ventilation), automated medical diagnosis from both structured and unstructured clinical data, and other domains such as health resource allocation, drug discovery, and health management. It argues that the RL framework—an agent learning optimal policies through trial-and-error interaction with an environment—maps naturally onto medical treatment as a sequential decision process. The authors assert that this survey is the first comprehensive survey of RL applications in healthcare.

Load-bearing premise

The claim that the survey is comprehensive and that its selected literature represents the field relies on an undocumented and potentially non-systematic search and inclusion process, rather than a transparent protocol.

Editorial extensions

If this is right

  • If RL-based treatment policies prove reliable in real clinical settings, they could enable personalized, adaptive treatment regimens that account for individual patient heterogeneity.
  • Successful integration of RL could reduce reliance on population-averaged treatment protocols by learning from accumulated patient data.
  • Addressing reward formulation challenges could lead to treatment policies that balance competing objectives such as efficacy and toxicity more explicitly.
  • Advances in off-policy evaluation and safe exploration would be needed before RL policies could be confidently deployed in clinical practice.
  • Small-data and transfer learning techniques could extend RL to rare diseases or new patient cohorts where historical data is scarce.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would be a systematic benchmark that evaluates offline RL treatment policies against standardized clinical outcomes across multiple ICU databases, going beyond the selective studies the survey reports.
  • The survey's emphasis on reward engineering suggests that progress may depend more on clinical expertise in specifying objectives than on advances in RL algorithms themselves.
  • The paper's catalog of applications implies that interoperable standards for state representation and reward specification could accelerate cross-institution validation of RL-based clinical decision support.
  • A consequence of the credit assignment discussion is that RL may be most easily validated in domains with short, well-defined decision horizons, such as medication dosing, before moving to longer-horizon interventions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper is a survey of reinforcement learning (RL) applications in healthcare. It begins with a review of RL foundations (MDPs, dynamic programming, Monte Carlo and temporal-difference methods, Q-learning, SARSA, policy search, actor-critic) and of key techniques (batch RL, model-based RL, transfer RL, representation learning for values/rewards/tasks, inverse RL, POMDPs). It then organizes the application literature into three broad areas: dynamic treatment regimes for chronic diseases and critical care, automated medical diagnosis from structured and unstructured data, and other healthcare domains such as resource allocation, process control, drug discovery, and health management. The final sections discuss challenges (state/action engineering, reward formulation, off-policy evaluation, model learning, exploration, credit assignment) and future directions (interpretability, prior knowledge, small data, ambient intelligence, and in-vivo validation). The conclusion states that the paper 'serves as the first comprehensive survey of RL applications in healthcare.'

Significance. If its coverage is accepted, the survey is a useful entry point for researchers new to RL in healthcare. Its strengths are the breadth of cited work, the clear application taxonomy, and the compact summary tables that map references to base methods, efficiency/representation techniques, data sources, and stated limitations. The treatment of off-policy evaluation, reward formulation, and the need for in-vivo validation is well aligned with the current state of the field. However, the paper's central contribution claim rests on 'first comprehensive survey,' and the manuscript provides no search protocol, inclusion/exclusion criteria, or comparison with prior surveys, so neither the priority nor the comprehensiveness can be independently verified. The paper also contains no limitation statement acknowledging this gap. These issues are load-bearing because comprehensive coverage is the stated basis of the survey's contribution, not merely a stylistic flourish.

major comments (2)
  1. [Section IX, with implications for Sections I and III-VI] The claim that the paper 'serves as the first comprehensive survey of RL applications in healthcare' is unsupported as written. The manuscript does not document a search protocol (databases, query terms, date range), inclusion/exclusion criteria, or any comparison with prior surveys, so a reader cannot check either 'first' or 'comprehensive.' Because this claim is the stated contribution of the paper, this is a load-bearing issue rather than a cosmetic one. I recommend either adding a methodology paragraph that makes the citation selection reproducible, or softening the claim to 'a survey' and explicitly stating the limitations of the coverage.
  2. [Section IX (conclusion) and the absence of a survey-limitations paragraph] The paper nowhere acknowledges the limitations of its own survey methodology. In particular, it does not state that the application map may be incomplete, that the selected references may be biased toward work visible to the authors, or that the catalog of challenges in Section VII is a synthesis of a non-systematically selected subset of the literature. Given that the paper explicitly builds its contribution on comprehensiveness, the absence of such a limitation statement should be corrected in the revision.
minor comments (5)
  1. [Section II-A, Eq. (7)] The epsilon-greedy policy in Eq. (7) is not a valid probability distribution: it assigns probability 1-epsilon to the greedy action and probability epsilon to every non-greedy action, which sums to more than 1 when the action set has more than two actions. The standard formulation is 1-epsilon+epsilon/|A| for the greedy action and epsilon/|A| for each non-greedy action.
  2. [Section IV-A1 and Table III] The acronym ODE is defined as 'Ordinary Difference Equations' in the text, but the cited models are systems of ordinary differential equations; the definition should read 'Ordinary Differential Equations.'
  3. [Section IV-A1, cancer chemotherapy paragraph] The text refers to 'support vector regression (SVG)' while Table III and the reference [104] describe support vector regression (SVR); the acronym should be corrected consistently to SVR.
  4. [Section II-A, Eq. (6)] The SARSA update in Eq. (6) is written as Q_t(s', pi(s')), but SARSA updates use the actually sampled next action a' drawn from the behavior policy; as written the notation conflates the behavior policy with the target policy and could confuse readers.
  5. [Section IV-B and Table IV] The text uses '3C (Compartimentalization, Corruption, and Complexity)'; the intended term appears to be 'Compartmentalization,' and the spelling should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this survey's taxonomy and challenge synthesis are independent of its inputs, and its authors' self-citations are descriptive rather than load-bearing.

full rationale

This paper is a survey, not a derivation. It does not claim to derive a novel result from first principles or to predict an outcome from fitted parameters. Its central content is an enumeration and organization of existing RL applications in healthcare, supported by external studies with their own methods, data, and evaluations. The organizational claims in Sections III-VI and the challenge lists in Sections VII-VIII are self-contained syntheses of the cited literature rather than reductions to the survey's own inputs. The authors' self-citations, such as [149] on causal policy gradient for HIV, [200] on deep IRL for sepsis treatment, [219] on inverse RL for mechanical ventilation, and [220] on supervised actor-critic for ventilation, appear as examples among many other works and are not used to justify the survey's conclusions. The 'first comprehensive survey' claim in Section IX is a completeness and historical claim; the absence of a documented search protocol is a coverage or methodological limitation, not a circularity defect. No fitted parameter is renamed as a prediction, no equation is defined in terms of its own output, and no load-bearing argument reduces to a self-citation. Therefore, per the default expectation for surveys, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

A survey introduces no free parameters, postulates, or entities. It relies on the accuracy and representativeness of the papers it cites. The only implicit premise is that the cited literature is a fair sample of the field, which is not established by a systematic search protocol.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Reinforcement Learning in Healthcare: A Survey." pith.science (2026). https://pith.science/paper/SOG4EKXG

@misc{pith2026190808796,
  author       = {Pith},
  title        = {Pith review of: Reinforcement Learning in Healthcare: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SOG4EKXG}},
  note         = {Machine review of arXiv:1908.08796}
}
read the original abstract

As a subfield of machine learning, reinforcement learning (RL) aims at empowering one's capabilities in behavioural decision making by using interaction experience with the world and an evaluative feedback. Unlike traditional supervised learning methods that usually rely on one-shot, exhaustive and supervised reward signals, RL tackles with sequential decision making problems with sampled, evaluative and delayed feedback simultaneously. Such distinctive features make RL technique a suitable candidate for developing powerful solutions in a variety of healthcare domains, where diagnosing decisions or treatment regimes are usually characterized by a prolonged and sequential procedure. This survey discusses the broad applications of RL techniques in healthcare domains, in order to provide the research community with systematic understanding of theoretical foundations, enabling methods and techniques, existing challenges, and new insights of this emerging paradigm. By first briefly examining theoretical foundations and key techniques in RL research from efficient and representational directions, we then provide an overview of RL applications in healthcare domains ranging from dynamic treatment regimes in chronic diseases and critical care, automated medical diagnosis from both unstructured and structured clinical data, as well as many other control or scheduling domains that have infiltrated many aspects of a healthcare system. Finally, we summarize the challenges and open issues in current research, and point out some potential solutions and directions for future research.

Figures

Figures reproduced from arXiv: 1908.08796 by the authors.

Figure 1
Figure 1. The summarization of theoretical foundations, basic solutions, challenging issues and advanced techniques in RL. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The outline of application domains of RL in healthcare. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Counterfactual Shapley Credit Assignment

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Counterfactual Shapley values, computed by simulated 'what-if' action replacements, redistribute RL rewards without changing the optimal policy and improve credit assignment in stochastic, sparse, delayed-reward tasks.

  2. Reinforcement Learning for a Discrete-Time Linear-Quadratic Control Problem with an Application

    stat.ML 2024-12 reject novelty 3.0 of 10

    The paper claims entropy regularization forces the optimal LQ feedback policy to be Gaussian and uses that to solve a mean-variance asset-liability problem, but the proof of the main theorem contains a correlation err...

Reference graph

Works this paper leans on

300 extracted references · 67 canonical work pages · cited by 2 Pith papers

  1. [1]

    The coming of age of artificial intelligence in medicine,

    V . L. Patel, E. H. Shortliffe, M. Stefanelli, P. Szolovits, M. R. Berthold, R. Bellazzi, and A. Abu-Hanna, “The coming of age of artificial intelligence in medicine,” Artificial Intelligence in Medicine , vol. 46, no. 1, pp. 5–17, 2009

  2. [2]

    Artificial intelligence in medicine and cardiac imaging: harnessing big data and advanced computing to provide personalized medical diagnosis and treatment,

    S. E. Dilsizian and E. L. Siegel, “Artificial intelligence in medicine and cardiac imaging: harnessing big data and advanced computing to provide personalized medical diagnosis and treatment,” Current Cardiology Reports, vol. 16, no. 1, p. 441, 2014

  3. [3]

    Artificial intelligence in healthcare: past, present and future,

    F. Jiang, Y . Jiang, H. Zhi, Y . Dong, H. Li, S. Ma, Y . Wang, Q. Dong, H. Shen, and Y . Wang, “Artificial intelligence in healthcare: past, present and future,” Stroke and Vascular Neurology, vol. 2, no. 4, pp. 230–243, 2017

  4. [4]

    The practical implementation of artificial intelligence technologies in medicine,

    J. He, S. L. Baxter, J. Xu, J. Xu, X. Zhou, and K. Zhang, “The practical implementation of artificial intelligence technologies in medicine,” Nature Medicine, vol. 25, no. 1, p. 30, 2019

  5. [5]

    Machine learning and decision support in critical care,

    A. E. Johnson, M. M. Ghassemi, S. Nemati, K. E. Niehaus, D. A. Clifton, and G. D. Clifford, “Machine learning and decision support in critical care,” Proceedings of the IEEE , vol. 104, no. 2, pp. 444–466, 2016

  6. [6]

    Deep learning for health informatics,

    D. Rav `ı, C. Wong, F. Deligianni, M. Berthelot, J. Andreu-Perez, B. Lo, and G.-Z. Yang, “Deep learning for health informatics,” IEEE Journal of Biomedical and Health Informatics , vol. 21, no. 1, pp. 4–21, 2017

  7. [7]

    Opportunities and obstacles for deep learning in biology and medicine,

    T. Ching, D. S. Himmelstein, B. K. Beaulieu-Jones, A. A. Kalinin, B. T. Do, G. P. Way, E. Ferrero, P.-M. Agapow, M. Zietz, M. M. Hoffman et al. , “Opportunities and obstacles for deep learning in biology and medicine,” bioRxiv, p. 142760, 2018

  8. [8]

    Big data application in biomedical research and health care: a literature review,

    J. Luo, M. Wu, D. Gopukumar, and Y . Zhao, “Big data application in biomedical research and health care: a literature review,” Biomedical Informatics Insights, vol. 8, pp. BII–S31 559, 2016

Show all 300 references
  1. [9]

    A guide to deep learning in healthcare,

    A. Esteva, A. Robicquet, B. Ramsundar, V . Kuleshov, M. DePristo, K. Chou, C. Cui, G. Corrado, S. Thrun, and J. Dean, “A guide to deep learning in healthcare,” Nature Medicine, vol. 25, no. 1, p. 24, 2019

  2. [10]

    Human-level control through deep reinforcement learning,

    V . Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, p. 529, 2015

  3. [11]

    Reinforcement learning improves behaviour from evaluative feedback,

    M. L. Littman, “Reinforcement learning improves behaviour from evaluative feedback,” Nature, vol. 521, no. 7553, p. 445, 2015

  4. [12]

    Deep reinforcement learning,

    Y . Li, “Deep reinforcement learning,”arXiv preprint arXiv:1810.06339, 2018

  5. [13]

    Applications of deep learning and reinforcement learning to biological data,

    M. Mahmud, M. S. Kaiser, A. Hussain, and S. Vassanelli, “Applications of deep learning and reinforcement learning to biological data,” IEEE transactions on neural networks and learning systems , vol. 29, no. 6, pp. 2063–2079, 2018

  6. [14]

    R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction. MIT press, 2018

  7. [15]

    Re- inforcement learning for control: Performance, stability, and deep approximators,

    L. Bus ¸oniu, T. de Bruin, D. Toli ´c, J. Kober, and I. Palunko, “Re- inforcement learning for control: Performance, stability, and deep approximators,” Annual Reviews in Control , 2018

  8. [16]

    Guidelines for reinforcement learning in healthcare

    O. Gottesman, F. Johansson, M. Komorowski, A. Faisal, D. Sontag, F. Doshi-Velez, and L. A. Celi, “Guidelines for reinforcement learning in healthcare.” Nature medicine, vol. 25, no. 1, p. 16, 2019

  9. [17]

    Bellman, Dynamic programming

    R. Bellman, Dynamic programming. Courier Corporation, 2013

  10. [18]

    Q-learning,

    C. J. Watkins and P. Dayan, “Q-learning,” Machine Learning, vol. 8, no. 3-4, pp. 279–292, 1992

  11. [19]

    G. A. Rummery and M. Niranjan, On-line Q-learning using connec- tionist systems. University of Cambridge, Department of Engineering Cambridge, England, 1994, vol. 37

  12. [20]

    Policy search for motor primitives in robotics,

    J. Kober and J. R. Peters, “Policy search for motor primitives in robotics,” inAdvances in Neural Information Processing Systems, 2009, pp. 849–856

  13. [21]

    Natural actor-critic,

    J. Peters and S. Schaal, “Natural actor-critic,” Neurocomputing, vol. 71, no. 7-9, pp. 1180–1190, 2008

  14. [22]

    Bayesian reinforcement learning,

    N. Vlassis, M. Ghavamzadeh, S. Mannor, and P. Poupart, “Bayesian reinforcement learning,” in Reinforcement Learning. Springer, 2012, pp. 359–386

  15. [23]

    Bayesian reinforcement learning: A survey,

    M. Ghavamzadeh, S. Mannor, J. Pineau, A. Tamar et al. , “Bayesian reinforcement learning: A survey,” Foundations and Trends R© in Ma- chine Learning, vol. 8, no. 5-6, pp. 359–483, 2015

  16. [24]

    Reinforcement learning in finite mdps: Pac analysis,

    A. L. Strehl, L. Li, and M. L. Littman, “Reinforcement learning in finite mdps: Pac analysis,” Journal of Machine Learning Research , vol. 10, no. Nov, pp. 2413–2444, 2009

  17. [25]

    R-max-a general polynomial time algorithm for near-optimal reinforcement learning,

    R. I. Brafman and M. Tennenholtz, “R-max-a general polynomial time algorithm for near-optimal reinforcement learning,”Journal of Machine Learning Research, vol. 3, no. Oct, pp. 213–231, 2002

  18. [26]

    Intrinsic motivation and reinforcement learning,

    A. G. Barto, “Intrinsic motivation and reinforcement learning,” in Intrinsically motivated learning in natural and artificial systems . Springer, 2013, pp. 17–47

  19. [27]

    A comprehensive survey of multiagent reinforcement learning,

    L. Busoniu, R. Babuska, and B. De Schutter, “A comprehensive survey of multiagent reinforcement learning,” IEEE Transactions on Systems, Man, And Cybernetics-Part C: Applications and Reviews, 38 (2), 2008 , 2008

  20. [28]

    On the sample complexity of reinforcement learning,

    S. M. Kakade et al. , “On the sample complexity of reinforcement learning,” Ph.D. dissertation, University of London London, England, 2003

  21. [29]

    Sample complexity bounds of exploration,

    L. Li, “Sample complexity bounds of exploration,” in Reinforcement Learning. Springer, 2012, pp. 175–204

  22. [30]

    Busoniu, R

    L. Busoniu, R. Babuska, B. De Schutter, and D. Ernst, Reinforcement learning and dynamic programming using function approximators . CRC press, 2010

  23. [31]

    Reinforcement learning in continuous state and action spaces,

    H. Van Hasselt, “Reinforcement learning in continuous state and action spaces,” in Reinforcement learning. Springer, 2012, pp. 207–251

  24. [32]

    Safe exploration in markov decision processes,

    T. M. Moldovan and P. Abbeel, “Safe exploration in markov decision processes,” in Proceedings of the 29th International Coference on International Conference on Machine Learning . Omnipress, 2012, pp. 1451–1458

  25. [33]

    A comprehensive survey on safe rein- forcement learning,

    J. Garcıa and F. Fern ´andez, “A comprehensive survey on safe rein- forcement learning,” Journal of Machine Learning Research , vol. 16, no. 1, pp. 1437–1480, 2015

  26. [34]

    Robust markov decision processes,

    W. Wiesemann, D. Kuhn, and B. Rustem, “Robust markov decision processes,” Mathematics of Operations Research , vol. 38, no. 1, pp. 153–183, 2013

  27. [35]

    Distributionally robust markov decision pro- cesses,

    H. Xu and S. Mannor, “Distributionally robust markov decision pro- cesses,” in Advances in Neural Information Processing Systems , 2010, pp. 2505–2513

  28. [36]

    Interpretable policies for rein- forcement learning by genetic programming,

    D. Hein, S. Udluft, and T. A. Runkler, “Interpretable policies for rein- forcement learning by genetic programming,”Engineering Applications of Artificial Intelligence , vol. 76, pp. 158–169, 2018

  29. [37]

    Verifiable reinforcement learning via policy extraction,

    O. Bastani, Y . Pu, and A. Solar-Lezama, “Verifiable reinforcement learning via policy extraction,” in Advances in Neural Information Processing Systems, 2018, pp. 2499–2509

  30. [38]

    Reinforcement learning,

    M. Wiering and M. Van Otterlo, “Reinforcement learning,” Adaptation, learning, and optimization , vol. 12, 2012

  31. [39]

    Batch reinforcement learning,

    S. Lange, T. Gabel, and M. Riedmiller, “Batch reinforcement learning,” in Reinforcement learning. Springer, 2012, pp. 45–73

  32. [40]

    Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method,

    M. Riedmiller, “Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method,” in European Confer- ence on Machine Learning . Springer, 2005, pp. 317–328

  33. [41]

    Tree-based batch mode rein- forcement learning,

    D. Ernst, P. Geurts, and L. Wehenkel, “Tree-based batch mode rein- forcement learning,” Journal of Machine Learning Research , vol. 6, no. Apr, pp. 503–556, 2005

  34. [42]

    Least-squares policy iteration,

    M. G. Lagoudakis and R. Parr, “Least-squares policy iteration,” Journal of Machine Learning Research , vol. 4, no. Dec, pp. 1107–1149, 2003

  35. [43]

    Learning and using models,

    T. Hester and P. Stone, “Learning and using models,” in Reinforcement learning. Springer, 2012, pp. 111–141

  36. [44]

    A survey of monte carlo tree search methods,

    C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton, “A survey of monte carlo tree search methods,” IEEE Transactions on Computational Intelligence and AI in games , vol. 4, no. 1, pp. 1–43, 2012

  37. [45]

    Transfer in reinforcement learning: a framework and a survey,

    A. Lazaric, “Transfer in reinforcement learning: a framework and a survey,” in Reinforcement Learning. Springer, 2012, pp. 143–173

  38. [46]

    Transfer learning for reinforcement learning domains: A survey,

    M. E. Taylor and P. Stone, “Transfer learning for reinforcement learning domains: A survey,” Journal of Machine Learning Research , vol. 10, no. Jul, pp. 1633–1685, 2009

  39. [47]

    Mastering the game of go with deep neural networks and tree search,

    D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V . Panneershelvam, M. Lanctot et al., “Mastering the game of go with deep neural networks and tree search,” Nature, vol. 529, no. 7587, p. 484, 2016

  40. [48]

    A survey of deep neural network architectures and their applications,

    W. Liu, Z. Wang, X. Liu, N. Zeng, Y . Liu, and F. E. Alsaadi, “A survey of deep neural network architectures and their applications,” Neurocomputing, vol. 234, pp. 11–26, 2017

  41. [49]

    Efficient processing of deep neural networks: A tutorial and survey,

    V . Sze, Y .-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient processing of deep neural networks: A tutorial and survey,” Proceedings of the IEEE, vol. 105, no. 12, pp. 2295–2329, 2017

  42. [50]

    Multiobjective reinforcement learning: A comprehensive overview,

    C. Liu, X. Xu, and D. Hu, “Multiobjective reinforcement learning: A comprehensive overview,” IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 45, no. 3, pp. 385–398, 2015

  43. [51]

    A survey of preference-based reinforcement learning methods,

    C. Wirth, R. Akrour, G. Neumann, and J. F ¨urnkranz, “A survey of preference-based reinforcement learning methods,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 4945–4990, 2017

  44. [52]

    Preference- based reinforcement learning: a formal framework and a policy iteration algorithm,

    J. F ¨urnkranz, E. H ¨ullermeier, W. Cheng, and S.-H. Park, “Preference- based reinforcement learning: a formal framework and a policy iteration algorithm,” Machine Learning, vol. 89, no. 1-2, pp. 123–156, 2012

  45. [53]

    Algorithms for inverse reinforcement learning

    A. Y . Ng, S. J. Russell et al. , “Algorithms for inverse reinforcement learning.” in ICML, 2000, pp. 663–670

  46. [54]

    A survey of inverse reinforcement learning techniques,

    S. Zhifei and E. Meng Joo, “A survey of inverse reinforcement learning techniques,” International Journal of Intelligent Computing and Cybernetics, vol. 5, no. 3, pp. 293–311, 2012

  47. [55]

    Maximum entropy inverse reinforcement learning

    B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning.” in AAAI, vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438

  48. [56]

    Apprenticeship learning via inverse rein- forcement learning,

    P. Abbeel and A. Y . Ng, “Apprenticeship learning via inverse rein- forcement learning,” in Proceedings of the twenty-first international conference on Machine learning . ACM, 2004, p. 1

  49. [57]

    Nonlinear inverse reinforcement learning with gaussian processes,

    S. Levine, Z. Popovic, and V . Koltun, “Nonlinear inverse reinforcement learning with gaussian processes,” in Advances in Neural Information Processing Systems, 2011, pp. 19–27

  50. [58]

    Bayesian inverse reinforcement learn- ing,

    D. Ramachandran and E. Amir, “Bayesian inverse reinforcement learn- ing,” Urbana, vol. 51, no. 61801, pp. 1–4, 2007

  51. [59]

    Efficient solution algorithms for factored mdps,

    C. Guestrin, D. Koller, R. Parr, and S. Venkataraman, “Efficient solution algorithms for factored mdps,”Journal of Artificial Intelligence Research, vol. 19, pp. 399–468, 2003

  52. [60]

    Efficient reinforcement learning in factored mdps,

    M. Kearns and D. Koller, “Efficient reinforcement learning in factored mdps,” in IJCAI, vol. 16, 1999, pp. 740–747

  53. [61]

    Algorithm-directed exploration for model-based reinforcement learning in factored mdps,

    C. Guestrin, R. Patrascu, and D. Schuurmans, “Algorithm-directed exploration for model-based reinforcement learning in factored mdps,” in ICML, 2002, pp. 235–242

  54. [62]

    Near-optimal reinforcement learning in factored mdps,

    I. Osband and B. Van Roy, “Near-optimal reinforcement learning in factored mdps,” inAdvances in Neural Information Processing Systems, 2014, pp. 604–612

  55. [63]

    Efficient structure learning in factored-state mdps,

    A. L. Strehl, C. Diuk, and M. L. Littman, “Efficient structure learning in factored-state mdps,” in AAAI, vol. 7, 2007, pp. 645–650

  56. [64]

    Recent advances in hierarchical reinforcement learning,

    A. G. Barto and S. Mahadevan, “Recent advances in hierarchical reinforcement learning,” Discrete Event Dynamic Systems , vol. 13, no. 1-2, pp. 41–77, 2003

  57. [65]

    Hierarchical approaches,

    B. Hengst, “Hierarchical approaches,” in Reinforcement learning . Springer, 2012, pp. 293–323

  58. [66]

    Solving relational and first-order logical markov decision processes: A survey,

    M. van Otterlo, “Solving relational and first-order logical markov decision processes: A survey,” in Reinforcement Learning. Springer, 2012, pp. 253–292

  59. [67]

    Reinforcement learning algorithm for partially observable markov decision problems,

    T. Jaakkola, S. P. Singh, and M. I. Jordan, “Reinforcement learning algorithm for partially observable markov decision problems,” in Ad- vances in Neural Information Processing Systems , 1995, pp. 345–352

  60. [68]

    A computer program for digitalis dosage regimens,

    R. W. Jelliffe, J. Buell, R. Kalaba, R. Sridhar, and R. Rockwell, “A computer program for digitalis dosage regimens,” Mathematical Biosciences, vol. 9, pp. 179–193, 1970

  61. [69]

    R. E. Bellman, Mathematical methods in medicine . World Scientific Publishing Co., Inc., 1983

  62. [70]

    Comparison of some control strategies for three-compartment pk/pd models,

    C. Hu, W. S. Lovejoy, and S. L. Shafer, “Comparison of some control strategies for three-compartment pk/pd models,” Journal of Pharmacokinetics and Biopharmaceutics , vol. 22, no. 6, pp. 525–550, 1994

  63. [71]

    Modeling medical treatment using markov decision processes,

    A. J. Schaefer, M. D. Bailey, S. M. Shechter, and M. S. Roberts, “Modeling medical treatment using markov decision processes,” in Operations Research and Health Care . Springer, 2005, pp. 593–612

  64. [72]

    Dynamic treatment regimes,

    B. Chakraborty and S. A. Murphy, “Dynamic treatment regimes,” Annual Review of Statistics and Its Application , vol. 1, pp. 447–464, 2014

  65. [73]

    Dynamic treatment regimes: Technical challenges and applications,

    E. B. Laber, D. J. Lizotte, M. Qian, W. E. Pelham, and S. A. Murphy, “Dynamic treatment regimes: Technical challenges and applications,” Electronic Journal of Statistics , vol. 8, no. 1, p. 1225, 2014

  66. [74]

    Estimation of survival distributions of treatment policies in two-stage randomization designs in clinical trials,

    J. K. Lunceford, M. Davidian, and A. A. Tsiatis, “Estimation of survival distributions of treatment policies in two-stage randomization designs in clinical trials,” Biometrics, vol. 58, no. 1, pp. 48–57, 2002

  67. [75]

    Adaptive interventions in child and adolescent mental health,

    D. Almirall and A. Chronis-Tuscano, “Adaptive interventions in child and adolescent mental health,” Journal of Clinical Child & Adolescent Psychology, vol. 45, no. 4, pp. 383–395, 2016

  68. [76]

    Adaptive treatment strategies in chronic disease,

    P. W. Lavori and R. Dawson, “Adaptive treatment strategies in chronic disease,” Annu. Rev. Med., vol. 59, pp. 443–453, 2008

  69. [77]

    Chakraborty and E

    B. Chakraborty and E. E. M. Moodie, Statistical Reinforcement Learn- ing. Springer New York, 2013

  70. [78]

    An experimental design for the development of adaptive treatment strategies,

    S. A. Murphy, “An experimental design for the development of adaptive treatment strategies,” Statistics in Medicine , vol. 24, no. 10, pp. 1455– 1481, 2005

  71. [79]

    Developing adaptive treatment strategies in substance abuse research,

    S. A. Murphy, K. G. Lynch, D. Oslin, J. R. McKay, and T. TenHave, “Developing adaptive treatment strategies in substance abuse research,” Drug & Alcohol Dependence , vol. 88, pp. S24–S30, 2007

  72. [80]

    W. H. Organization, Preventing chronic diseases: a vital investment . World Health Organization, 2005

  73. [81]

    Chakraborty and E

    B. Chakraborty and E. Moodie, Statistical methods for dynamic treat- ment regimes. Springer, 2013

  74. [82]

    Improving chronic illness care: translating evidence into action,

    E. H. Wagner, B. T. Austin, C. Davis, M. Hindmarsh, J. Schaefer, and A. Bonomi, “Improving chronic illness care: translating evidence into action,” Health Affairs, vol. 20, no. 6, pp. 64–78, 2001

  75. [83]

    Reinforcement learning design for cancer clinical trials,

    Y . Zhao, M. R. Kosorok, and D. Zeng, “Reinforcement learning design for cancer clinical trials,” Statistics in Medicine , vol. 28, no. 26, pp. 3294–3315, 2009

  76. [84]

    Reinforcement learning based control of tumor growth with chemotherapy,

    A. Hassani et al. , “Reinforcement learning based control of tumor growth with chemotherapy,” in 2010 International Conference on System Science and Engineering (ICSSE) . IEEE, 2010, pp. 185–189

  77. [85]

    Drug scheduling of cancer chemotherapy based on natural actor-critic approach,

    I. Ahn and J. Park, “Drug scheduling of cancer chemotherapy based on natural actor-critic approach,” BioSystems, vol. 106, no. 2-3, pp. 121–129, 2011

  78. [86]

    Using reinforcement learning to personalize dosing strategies in a simulated cancer trial with high dimensional data,

    K. Humphrey, “Using reinforcement learning to personalize dosing strategies in a simulated cancer trial with high dimensional data,” 2017

  79. [87]

    Reinforcement learning-based control of drug dosing for cancer chemotherapy treat- ment,

    R. Padmanabhan, N. Meskin, and W. M. Haddad, “Reinforcement learning-based control of drug dosing for cancer chemotherapy treat- ment,” Mathematical biosciences, vol. 293, pp. 11–20, 2017

  80. [88]

    Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer,

    Y . Zhao, D. Zeng, M. A. Socinski, and M. R. Kosorok, “Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer,” Biometrics, vol. 67, no. 4, pp. 1422–1433, 2011

  81. [89]

    Preference- based policy iteration: Leveraging preference learning for reinforce- ment learning,

    W. Cheng, J. F ¨urnkranz, E. H ¨ullermeier, and S.-H. Park, “Preference- based policy iteration: Leveraging preference learning for reinforce- ment learning,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2011, pp. 312– 327

  82. [90]

    April: Active preference learning-based reinforcement learning,

    R. Akrour, M. Schoenauer, and M. Sebag, “April: Active preference learning-based reinforcement learning,” in Joint European Confer- ence on Machine Learning and Knowledge Discovery in Databases . Springer, 2012, pp. 116–131

  83. [91]

    Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm,

    R. Busa-Fekete, B. Sz ¨or´enyi, P. Weng, W. Cheng, and E. H ¨ullermeier, “Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm,” Machine Learning, vol. 97, no. 3, pp. 327–351, 2014

  84. [92]

    Reinforcement learning in models of adaptive medical treatment strategies,

    R. Vincent, “Reinforcement learning in models of adaptive medical treatment strategies,” Ph.D. dissertation, McGill University Libraries, 2014

  85. [93]

    Deep reinforcement learning for automated radiation adaptation in lung cancer,

    H. H. Tseng, Y . Luo, S. Cui, J. T. Chien, R. K. Ten Haken, and I. E. Naqa, “Deep reinforcement learning for automated radiation adaptation in lung cancer,”Medical Physics, vol. 44, no. 12, pp. 6690–6705, 2017

  86. [94]

    Simulation-based optimization of radiotherapy: Agent-based modeling and reinforcement learning,

    A. Jalalimanesh, H. S. Haghighi, A. Ahmadi, and M. Soltani, “Simulation-based optimization of radiotherapy: Agent-based modeling and reinforcement learning,” Mathematics and Computers in Simula- tion, vol. 133, pp. 235–248, 2017

  87. [95]

    Multi-objective optimization of radiotherapy: distributed q-learning and agent-based simulation,

    A. Jalalimanesh, H. S. Haghighi, A. Ahmadi, H. Hejazian, and M. Soltani, “Multi-objective optimization of radiotherapy: distributed q-learning and agent-based simulation,” Journal of Experimental & Theoretical Artificial Intelligence , pp. 1–16, 2017

  88. [96]

    Q-learning with censored data,

    Y . Goldberg and M. R. Kosorok, “Q-learning with censored data,” Annals of Statistics , vol. 40, no. 1, p. 529, 2012

  89. [97]

    Personalized medical treatments using novel reinforce- ment learning algorithms,

    Y . M. Soliman, “Personalized medical treatments using novel reinforce- ment learning algorithms,” arXiv preprint arXiv:1406.3922 , 2014

  90. [98]

    Reinforcement learning with action-derived rewards for chemotherapy and clinical trial dosing regimen selection,

    G. Yauney and P. Shah, “Reinforcement learning with action-derived rewards for chemotherapy and clinical trial dosing regimen selection,” in Machine Learning for Healthcare Conference , 2018, pp. 161–226

  91. [99]

    World cancer report 2014,

    B. Stewart, C. P. Wild et al., “World cancer report 2014,”Health, 2017

  92. [100]

    Interactions between the immune system and cancer: a brief review of non-spatial mathematical models,

    R. Eftimie, J. L. Bramson, and D. J. Earn, “Interactions between the immune system and cancer: a brief review of non-spatial mathematical models,” Bulletin of Mathematical Biology , vol. 73, no. 1, pp. 2–32, 2011

  93. [101]

    A survey of optimiza- tion models on cancer chemotherapy treatment planning,

    J. Shi, O. Alagoz, F. S. Erenay, and Q. Su, “A survey of optimiza- tion models on cancer chemotherapy treatment planning,” Annals of Operations Research, vol. 221, no. 1, pp. 331–356, 2014

  94. [102]

    Cancer evolution: mathematical models and computational inference,

    N. Beerenwinkel, R. F. Schwarz, M. Gerstung, and F. Markowetz, “Cancer evolution: mathematical models and computational inference,” Systematic Biology, vol. 64, no. 1, pp. e1–e25, 2014

  95. [103]

    Person- alizing cancer therapy via machine learning,

    M. Tenenbaum, A. Fern, L. Getoor, M. Littman, V . Manasinghka, S. Natarajan, D. Page, J. Shrager, Y . Singer, and P. Tadepalli, “Person- alizing cancer therapy via machine learning,” in Workshops of NIPS , 2010

  96. [104]

    Support vector method for function approximation, regression estimation and signal processing,

    V . Vapnik, S. E. Golowich, and A. J. Smola, “Support vector method for function approximation, regression estimation and signal processing,” in Advances in Neural Information Processing Systems, 1997, pp. 281– 287

  97. [105]

    The dynamics of an optimally controlled tumor model: A case study,

    L. G. De Pillis and A. Radunskaya, “The dynamics of an optimally controlled tumor model: A case study,” Mathematical and Computer Modelling, vol. 37, no. 11, pp. 1221–1244, 2003

  98. [106]

    Machine learning in radiation oncology: Opportunities, requirements, and needs,

    M. Feng, G. Valdes, N. Dixit, and T. D. Solberg, “Machine learning in radiation oncology: Opportunities, requirements, and needs,” Frontiers in Oncology, vol. 8, 2018

  99. [107]

    Robust high performance reinforce- ment learning through weighted k-nearest neighbors,

    J. de Lope, D. Maravall et al. , “Robust high performance reinforce- ment learning through weighted k-nearest neighbors,”Neurocomputing, vol. 74, no. 8, pp. 1251–1259, 2011

  100. [108]

    Idf diabetes atlas: Global estimates of diabetes prevalence for 2017 and projections for 2045,

    N. Cho, J. Shaw, S. Karuranga, Y . Huang, J. da Rocha Fernandes, A. Ohlrogge, and B. Malanda, “Idf diabetes atlas: Global estimates of diabetes prevalence for 2017 and projections for 2045,” Diabetes Research and Clinical Practice , vol. 138, pp. 271–281, 2018

  101. [109]

    Clinical control of diabetes by the artificial pancreas,

    A. M. Albisser, B. Leibel, T. Ewart, Z. Davidovac, C. Botz, W. Zingg, H. Schipper, and R. Gander, “Clinical control of diabetes by the artificial pancreas,” Diabetes, vol. 23, no. 5, pp. 397–404, 1974

  102. [110]

    Artificial pancreas: past, present, future,

    C. Cobelli, E. Renard, and B. Kovatchev, “Artificial pancreas: past, present, future,” Diabetes, vol. 60, no. 11, pp. 2672–2682, 2011

  103. [111]

    A critical assessment of algorithms and challenges in the development of a closed-loop artificial pancreas,

    B. W. Bequette, “A critical assessment of algorithms and challenges in the development of a closed-loop artificial pancreas,” Diabetes Technology & Therapeutics, vol. 7, no. 1, pp. 28–47, 2005

  104. [112]

    The artificial pancreas: current status and future prospects in the management of diabetes,

    T. Peyser, E. Dassau, M. Breton, and J. S. Skyler, “The artificial pancreas: current status and future prospects in the management of diabetes,” Annals of the New York Academy of Sciences , vol. 1311, no. 1, pp. 102–123, 2014

  105. [113]

    The use of reinforcement learning algorithms to meet the challenges of an artificial pancreas,

    M. K. Bothe, L. Dickens, K. Reichel, A. Tellmann, B. Ellger, M. West- phal, and A. A. Faisal, “The use of reinforcement learning algorithms to meet the challenges of an artificial pancreas,” Expert Review of Medical Devices, vol. 10, no. 5, pp. 661–673, 2013

  106. [114]

    Agent-based sim- ulation for blood glucose,

    S. Yasini, M. B. Naghibi Sistani, and A. Karimpour, “Agent-based sim- ulation for blood glucose,” International Journal of Applied Science, Engineering and Technology, vol. 5, pp. 89–95, 2009

  107. [115]

    Pre- liminary results of a novel approach for glucose regulation using an actor-critic learning based controller,

    E. Daskalaki, L. Scarnato, P. Diem, and S. G. Mougiakakou, “Pre- liminary results of a novel approach for glucose regulation using an actor-critic learning based controller,” 2010

  108. [116]

    In silico preclinical trials: a proof of concept in closed-loop control of type 1 diabetes,

    B. P. Kovatchev, M. Breton, C. Dalla Man, and C. Cobelli, “In silico preclinical trials: a proof of concept in closed-loop control of type 1 diabetes,” 2009

  109. [117]

    An actor–critic based controller for glucose regulation in type 1 diabetes,

    E. Daskalaki, P. Diem, and S. G. Mougiakakou, “An actor–critic based controller for glucose regulation in type 1 diabetes,”Computer Methods and Programs in Biomedicine , vol. 109, no. 2, pp. 116–125, 2013

  110. [118]

    Personalized tuning of a reinforcement learning control al- gorithm for glucose regulation,

    ——, “Personalized tuning of a reinforcement learning control al- gorithm for glucose regulation,” in 2013 35th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). IEEE, 2013, pp. 3487–3490

  111. [119]

    Model-free machine learning in biomedicine: Feasibility study in type 1 diabetes,

    ——, “Model-free machine learning in biomedicine: Feasibility study in type 1 diabetes,” PloS One, vol. 11, no. 7, p. e0158722, 2016

  112. [120]

    A dual mode adaptive basal-bolus advisor based on reinforcement learning,

    Q. Sun, M. Jankovic, J. Budzinski, B. Moore, P. Diem, C. Stettler, and S. G. Mougiakakou, “A dual mode adaptive basal-bolus advisor based on reinforcement learning,” IEEE journal of biomedical and health informatics, 2018

  113. [121]

    Qualitative behavior of a family of delay-differential models of the glucose-insulin system,

    P. Palumbo, S. Panunzi, and A. De Gaetano, “Qualitative behavior of a family of delay-differential models of the glucose-insulin system,” Discrete and Continuous Dynamical Systems Series B , vol. 7, no. 2, p. 399, 2007

  114. [122]

    Glucose level control using temporal difference methods,

    A. Noori, M. A. Sadrnia et al., “Glucose level control using temporal difference methods,” in 2017 Iranian Conference on Electrical Engi- neering (ICEE). IEEE, 2017, pp. 895–900

  115. [123]

    Reinforcement-learning optimal control for type-1 diabetes,

    P. D. Ngo, S. Wei, A. Holubov ´a, J. Muzik, and F. Godtliebsen, “Reinforcement-learning optimal control for type-1 diabetes,” in 2018 IEEE EMBS International Conference on Biomedical & Health Infor- matics (BHI). IEEE, 2018, pp. 333–336

  116. [124]

    Control of blood glucose for type-1 diabetes by using rein- forcement learning with feedforward algorithm,

    ——, “Control of blood glucose for type-1 diabetes by using rein- forcement learning with feedforward algorithm,” Computational and Mathematical Methods in Medicine , vol. 2018, 2018

  117. [125]

    Quantitative estimation of insulin sensitivity

    R. N. Bergman, Y . Z. Ider, C. R. Bowden, and C. Cobelli, “Quantitative estimation of insulin sensitivity.” American Journal of Physiology- Endocrinology And Metabolism , vol. 236, no. 6, p. E667, 1979

  118. [126]

    Nonlinear model predictive control of glucose concen- tration in subjects with type 1 diabetes,

    R. Hovorka, V . Canonico, L. J. Chassin, U. Haueter, M. Massi- Benedetti, M. O. Federici, T. R. Pieber, H. C. Schaller, L. Schaupp, T. Veringet al., “Nonlinear model predictive control of glucose concen- tration in subjects with type 1 diabetes,” Physiological Measurement, vol...

  119. [127]

    Controlling blood glucose variability under uncertainty using reinforcement learning and gaussian processes,

    M. De Paula, L. O. ´Avila, and E. C. Mart ´ınez, “Controlling blood glucose variability under uncertainty using reinforcement learning and gaussian processes,” Applied Soft Computing , vol. 35, pp. 310–332, 2015

  120. [128]

    On-line policy learning and adaptation for real-time personalization of an artificial pancreas,

    M. De Paula, G. G. Acosta, and E. C. Mart ´ınez, “On-line policy learning and adaptation for real-time personalization of an artificial pancreas,” Expert Systems with Applications , vol. 42, no. 4, pp. 2234– 2255, 2015

  121. [129]

    Blood glucose regulation with stochastic optimal control for insulin-dependent diabetic patients,

    S. U. Acikgoz and U. M. Diwekar, “Blood glucose regulation with stochastic optimal control for insulin-dependent diabetic patients,” Chemical Engineering Science , vol. 65, no. 3, pp. 1227–1236, 2010

  122. [130]

    Modeling medical records of diabetes using markov decision processes,

    H. Asoh, M. Shiro, S. Akaho, T. Kamishima, K. Hashida, E. Aramaki, and T. Kohro, “Modeling medical records of diabetes using markov decision processes,” in Proceedings of ICML2013 Workshop on Role of Machine Learning in Transforming Healthcare , 2013

  123. [131]

    An application of inverse reinforcement learning to medical records of diabetes treatment,

    H. Asoh, M. S. S. Akaho, T. Kamishima, K. Hasida, E. Aramaki, and T. Kohro, “An application of inverse reinforcement learning to medical records of diabetes treatment,” in ECMLPKDD2013 Workshop on Reinforcement Learning with Generalized Feedback , 2013

  124. [132]

    Estimating dynamic treatment regimes in mobile health using v-learning,

    D. J. Luckett, E. B. Laber, A. R. Kahkoska, D. M. Maahs, E. Mayer- Davis, and M. R. Kosorok, “Estimating dynamic treatment regimes in mobile health using v-learning,” Journal of the American Statistical Association, no. just-accepted, pp. 1–39, 2018

  125. [133]

    Reinforcement learning approach to individualization of chronic pharmacotherapy,

    A. E. Gaweda, M. K. Muezzinoglu, G. R. Aronoff, A. A. Jacobs, J. M. Zurada, and M. E. Brier, “Reinforcement learning approach to individualization of chronic pharmacotherapy,” in IJCNN’05, vol. 5. IEEE, 2005, pp. 3290–3295

  126. [134]

    Model predictive control with reinforcement learning for drug delivery in renal anemia management,

    A. E. Gaweda, M. K. Muezzinoglu, A. A. Jacobs, G. R. Aronoff, and M. E. Brier, “Model predictive control with reinforcement learning for drug delivery in renal anemia management,” in IEEE EMBS’06. IEEE, 2006, pp. 5177–5180

  127. [135]

    Individualization of pharmacological anemia management using reinforcement learning,

    A. E. Gaweda, M. K. Muezzinoglu, G. R. Aronoff, A. A. Jacobs, J. M. Zurada, and M. E. Brier, “Individualization of pharmacological anemia management using reinforcement learning,” Neural Networks, vol. 18, no. 5-6, pp. 826–834, 2005

  128. [136]

    A reinforcement learn- ing approach for individualizing erythropoietin dosages in hemodialysis patients,

    J. D. Mart ´ın-Guerrero, F. Gomez, E. Soria-Olivas, J. Schmidhuber, M. Climente-Mart´ı, and N. V . Jim´enez-Torres, “A reinforcement learn- ing approach for individualizing erythropoietin dosages in hemodialysis patients,” Expert Systems with Applications , vol. 36, no. 6, pp....

  129. [137]

    Validation of a reinforcement learning policy for dosage optimization of erythropoietin,

    J. D. Mart ´ın-Guerrero, E. Soria-Olivas, M. Mart ´ınez-Sober, M. Climente-Mart ´ı, T. De Diego-Santos, and N. V . Jim ´enez- Torres, “Validation of a reinforcement learning policy for dosage optimization of erythropoietin,” in Australasian Joint Conference on Artificial Intell...

  130. [138]

    Optimizing drug therapy with reinforcement learning: The case of anemia management,

    J. M. Malof and A. E. Gaweda, “Optimizing drug therapy with reinforcement learning: The case of anemia management,” in Neural Networks (IJCNN), The 2011 International Joint Conference on. IEEE, 2011, pp. 2088–2092

  131. [139]

    Adaptive treatment of anemia on hemodialysis patients: A reinforcement learning approach,

    P. Escandell-Montero, J. M. Mart ´ınez-Mart´ınez, J. D. Mart´ın-Guerrero, E. Soria-Olivas, J. Vila-Franc´es, and R. Magdalena-Benedito, “Adaptive treatment of anemia on hemodialysis patients: A reinforcement learning approach,” in CIDM2011. IEEE, 2011, pp. 44–49

  132. [140]

    Optimization of anemia treatment in hemodialysis patients via reinforcement learning,

    P. Escandell-Montero, M. Chermisi, J. M. Martinez-Martinez, J. Gomez-Sanchis, C. Barbieri, E. Soria-Olivas, F. Mari, J. Vila- Franc´es, A. Stopper, E. Gatti et al., “Optimization of anemia treatment in hemodialysis patients via reinforcement learning,” Artificial Intelli- gence...

  133. [141]

    Dynamic multidrug therapies for hiv: Optimal and sti control approaches,

    B. M. Adams, H. T. Banks, H.-D. Kwon, and H. T. Tran, “Dynamic multidrug therapies for hiv: Optimal and sti control approaches,” Mathematical Biosciences and Engineering, vol. 1, no. 2, pp. 223–241, 2004

  134. [142]

    Clinical data based optimal sti strategies for hiv: a reinforcement learning approach,

    D. Ernst, G.-B. Stan, J. Goncalves, and L. Wehenkel, “Clinical data based optimal sti strategies for hiv: a reinforcement learning approach,” in 45th IEEE Conference on Decision and Control . IEEE, 2006, pp. 667–672

  135. [143]

    A reinforcement learning design for hiv clinical trials,

    S. Parbhoo, “A reinforcement learning design for hiv clinical trials,” Ph.D. dissertation, 2014

  136. [144]

    Combining kernel and model based learning for hiv therapy selection,

    S. Parbhoo, J. Bogojeska, M. Zazzi, V . Roth, and F. Doshi-Velez, “Combining kernel and model based learning for hiv therapy selection,” AMIA Summits on Translational Science Proceedings , vol. 2017, p. 239, 2017

  137. [145]

    Quanti- fying uncertainty in batch personalized sequential decision making

    V . N. Marivate, J. Chemali, E. Brunskill, and M. L. Littman, “Quanti- fying uncertainty in batch personalized sequential decision making.” in AAAI Workshop: Modern Artificial Intelligence for Health Analytics , 2014

  138. [146]

    Transfer learning across patient variations with hidden parameter markov decision processes,

    T. Killian, G. Konidaris, and F. Doshi-Velez, “Transfer learning across patient variations with hidden parameter markov decision processes,” arXiv preprint arXiv:1612.00475 , 2016

  139. [147]

    Robust and efficient transfer learning with hidden parameter markov decision processes,

    T. W. Killian, S. Daulton, G. Konidaris, and F. Doshi-Velez, “Robust and efficient transfer learning with hidden parameter markov decision processes,” in Advances in Neural Information Processing Systems , 2017, pp. 6250–6261

  140. [148]

    Direct policy transfer via hidden parameter markov decision processes,

    J. Yao, T. Killian, G. Konidaris, and F. Doshi-Velez, “Direct policy transfer via hidden parameter markov decision processes,” 2018

  141. [149]

    Incorporating causal factors into reinforcement learning for dynamic treatment regimes in hiv,

    C. Yu, Y . Dong, J. Liu, and G. Ren, “Incorporating causal factors into reinforcement learning for dynamic treatment regimes in hiv,” BMC medical informatics and decision making , vol. 19, no. 2, p. 60, 2019

  142. [150]

    Pac optimal exploration in continuous space markov decision processes

    J. Pazis and R. Parr, “Pac optimal exploration in continuous space markov decision processes.” in AAAI, 2013

  143. [151]

    Bounded optimal exploration in mdp

    K. Kawaguchi, “Bounded optimal exploration in mdp.” in AAAI, 2016, pp. 1758–1764

  144. [152]

    Methodological challenges in constructing effective treatment sequences for chronic psychiatric disorders,

    S. A. Murphy, D. W. Oslin, A. J. Rush, and J. Zhu, “Methodological challenges in constructing effective treatment sequences for chronic psychiatric disorders,” Neuropsychopharmacology, vol. 32, no. 2, p. 257, 2007

  145. [153]

    Eeg seizure detection and prediction algorithms: a survey,

    T. N. Alotaiby, S. A. Alshebeili, T. Alshawi, I. Ahmad, and F. E. A. El-Samie, “Eeg seizure detection and prediction algorithms: a survey,” EURASIP Journal on Advances in Signal Processing , vol. 2014, no. 1, p. 183, 2014

  146. [154]

    Progress in neuroengineering for brain repair: New challenges and open issues,

    G. Panuccio, M. Semprini, L. Natale, S. Buccelli, I. Colombi, and M. Chiappalone, “Progress in neuroengineering for brain repair: New challenges and open issues,” Brain and Neuroscience Advances, vol. 2, p. 2398212818776475, 2018

  147. [155]

    Adaptive treatment of epilepsy via batch-mode reinforcement learning

    A. Guez, R. D. Vincent, M. Avoli, and J. Pineau, “Adaptive treatment of epilepsy via batch-mode reinforcement learning.” in AAAI, 2008, pp. 1671–1678

  148. [156]

    Treating epilepsy via adaptive neurostimulation: a reinforcement learning ap- proach,

    J. Pineau, A. Guez, R. Vincent, G. Panuccio, and M. Avoli, “Treating epilepsy via adaptive neurostimulation: a reinforcement learning ap- proach,” International Journal of Neural Systems , vol. 19, no. 04, pp. 227–240, 2009

  149. [157]

    Adaptive control of epileptic seizures using reinforcement learning,

    A. Guez, “Adaptive control of epileptic seizures using reinforcement learning,” Ph.D. dissertation, McGill University Library, 2010

  150. [158]

    Adaptive control of epileptiform excitability in an in vitro model of limbic seizures,

    G. Panuccio, A. Guez, R. Vincent, M. Avoli, and J. Pineau, “Adaptive control of epileptiform excitability in an in vitro model of limbic seizures,” Experimental Neurology, vol. 241, pp. 179–183, 2013

  151. [159]

    Manifold embeddings for model-based rein- forcement learning under partial observability,

    K. Bush and J. Pineau, “Manifold embeddings for model-based rein- forcement learning under partial observability,” in Advances in Neural Information Processing Systems , 2009, pp. 189–197

  152. [160]

    Seizure control in a computational model using a reinforcement learning stimulation paradigm,

    V . Nagaraj, A. Lamperski, and T. I. Netoff, “Seizure control in a computational model using a reinforcement learning stimulation paradigm,” International Journal of Neural Systems , vol. 27, no. 07, p. 1750012, 2017

  153. [161]

    Sequenced treatment alternatives to relieve depression (star* d): rationale and design,

    A. J. Rush, M. Fava, S. R. Wisniewski, P. W. Lavori, M. H. Trivedi, H. A. Sackeim, M. E. Thase, A. A. Nierenberg, F. M. Quitkin, T. M. Kashner et al., “Sequenced treatment alternatives to relieve depression (star* d): rationale and design,”Controlled clinical trials, vol. 25, ...

  154. [162]

    Constructing evidence-based treatment strategies using methods from computer science,

    J. Pineau, M. G. Bellemare, A. J. Rush, A. Ghizaru, and S. A. Murphy, “Constructing evidence-based treatment strategies using methods from computer science,” Drug & Alcohol Dependence, vol. 88, pp. S52–S60, 2007

  155. [163]

    Kernel-based reinforcement learning,

    D. Ormoneit and ´S. Sen, “Kernel-based reinforcement learning,” Ma- chine learning, vol. 49, no. 2-3, pp. 161–178, 2002

  156. [164]

    Inference for optimal dynamic treatment regimes using an adaptive m-out-of-n bootstrap scheme,

    B. Chakraborty, E. B. Laber, and Y . Zhao, “Inference for optimal dynamic treatment regimes using an adaptive m-out-of-n bootstrap scheme,” Biometrics, vol. 69, no. 3, pp. 714–723, 2013

  157. [165]

    Interactive model building for q-learning,

    E. B. Laber, K. A. Linn, and L. A. Stefanski, “Interactive model building for q-learning,” Biometrika, vol. 101, no. 4, pp. 831–847, 2014

  158. [166]

    Interactive q-learning for probabilities and quantiles,

    K. A. Linn, E. B. Laber, and L. A. Stefanski, “Interactive q-learning for probabilities and quantiles,” arXiv preprint arXiv:1407.3414, 2014

  159. [167]

    Interactive q-learning for quantiles,

    ——, “Interactive q-learning for quantiles,” Journal of the American Statistical Association, vol. 112, no. 518, pp. 638–649, 2017

  160. [168]

    Q-and a- learning methods for estimating optimal dynamic treatment regimes,

    P. J. Schulte, A. A. Tsiatis, E. B. Laber, and M. Davidian, “Q-and a- learning methods for estimating optimal dynamic treatment regimes,” Statistical science: a review journal of the Institute of Mathematical Statistics, vol. 29, no. 4, p. 640, 2014

  161. [169]

    Optimal dynamic treatment regimes,

    S. A. Murphy, “Optimal dynamic treatment regimes,” Journal of the Royal Statistical Society: Series B (Statistical Methodology) , vol. 65, no. 2, pp. 331–355, 2003

  162. [170]

    Penalized q-learning for dynamic treatment regimens,

    R. Song, W. Wang, D. Zeng, and M. R. Kosorok, “Penalized q-learning for dynamic treatment regimens,” Statistica Sinica , vol. 25, no. 3, p. 901, 2015

  163. [171]

    Robust hybrid learning for estimating personalized dynamic treatment regimens,

    Y . Liu, Y . Wang, M. R. Kosorok, Y . Zhao, and D. Zeng, “Robust hybrid learning for estimating personalized dynamic treatment regimens,” arXiv preprint arXiv:1611.02314 , 2016

  164. [172]

    Budgeted learning for developing personalized treatment,

    K. Deng, R. Greiner, and S. Murphy, “Budgeted learning for developing personalized treatment,” in ICMLA2014. IEEE, 2014, pp. 7–14

  165. [173]

    Neurocognitive effects of antipsychotic medications in patients with chronic schizophrenia in the catie trial,

    R. S. Keefe, R. M. Bilder, S. M. Davis, P. D. Harvey, B. W. Palmer, J. M. Gold, H. Y . Meltzer, M. F. Green, G. Capuano, T. S. Stroup et al., “Neurocognitive effects of antipsychotic medications in patients with chronic schizophrenia in the catie trial,” Archives of General Ps...

  166. [174]

    Informing sequential clinical decision-making through reinforcement learning: an empirical study,

    S. M. Shortreed, E. Laber, D. J. Lizotte, T. S. Stroup, J. Pineau, and S. A. Murphy, “Informing sequential clinical decision-making through reinforcement learning: an empirical study,”Machine Learning, vol. 84, no. 1-2, pp. 109–136, 2011

  167. [175]

    Q-learning residual analysis: application to the effectiveness of sequences of antipsychotic medications for patients with schizophrenia,

    A. Ertefaie, S. Shortreed, and B. Chakraborty, “Q-learning residual analysis: application to the effectiveness of sequences of antipsychotic medications for patients with schizophrenia,” Statistics in Medicine , vol. 35, no. 13, pp. 2221–2234, 2016

  168. [176]

    Linear fitted-q iter- ation with multiple reward functions,

    D. J. Lizotte, M. Bowling, and S. A. Murphy, “Linear fitted-q iter- ation with multiple reward functions,” Journal of Machine Learning Research, vol. 13, no. Nov, pp. 3253–3295, 2012

  169. [177]

    Multi-objective markov decision processes for data-driven decision support,

    D. J. Lizotte and E. B. Laber, “Multi-objective markov decision processes for data-driven decision support,” The Journal of Machine Learning Research, vol. 17, no. 1, pp. 7378–7405, 2016

  170. [178]

    Set-valued dynamic treatment regimes for competing outcomes,

    E. B. Laber, D. J. Lizotte, and B. Ferguson, “Set-valued dynamic treatment regimes for competing outcomes,” Biometrics, vol. 70, no. 1, pp. 53–61, 2014

  171. [179]

    Incor- porating patient preferences into estimation of optimal individualized treatment rules,

    E. L. Butler, E. B. Laber, S. M. Davis, and M. R. Kosorok, “Incor- porating patient preferences into estimation of optimal individualized treatment rules,” Biometrics, 2017

  172. [180]

    Managing addiction as a chronic con- dition,

    M. Dennis and C. K. Scott, “Managing addiction as a chronic con- dition,” Addiction Science & Clinical Practice , vol. 4, no. 1, p. 45, 2007

  173. [181]

    A batch, off-policy, actor-critic algorithm for optimizing the average reward,

    S. A. Murphy, Y . Deng, E. B. Laber, H. R. Maei, R. S. Sutton, and K. Witkiewitz, “A batch, off-policy, actor-critic algorithm for optimizing the average reward,” arXiv preprint arXiv:1607.05047 , 2016

  174. [182]

    Inference for non-regular parameters in optimal dynamic treatment regimes,

    B. Chakraborty, S. Murphy, and V . Strecher, “Inference for non-regular parameters in optimal dynamic treatment regimes,” Statistical Methods in Medical Research , vol. 19, no. 3, pp. 317–343, 2010

  175. [183]

    Bias correction and confidence intervals for fitted q-iteration,

    B. Chakraborty, V . Strecher, and S. Murphy, “Bias correction and confidence intervals for fitted q-iteration,” in Workshop on Model Uncertainty and Risk in Reinforcement Learning, NIPS, Whistler, Canada. Citeseer, 2008

  176. [184]

    Tree-based reinforcement learning for estimating optimal dynamic treatment regimes,

    Y . Tao, L. Wang, D. Almirall et al., “Tree-based reinforcement learning for estimating optimal dynamic treatment regimes,” The Annals of Applied Statistics, vol. 12, no. 3, pp. 1914–1938, 2018

  177. [185]

    Critical care-where have we been and where are we going?

    J.-L. Vincent, “Critical care-where have we been and where are we going?” Critical Care, vol. 17, no. 1, p. S2, 2013

  178. [186]

    Critical care workforce,

    K. Krell, “Critical care workforce,” Critical Care Medicine , vol. 36, no. 4, pp. 1350–1353, 2008

  179. [187]

    State of the art review: the data revolution in critical care,

    M. Ghassemi, L. A. Celi, and D. J. Stone, “State of the art review: the data revolution in critical care,” Critical Care, vol. 19, no. 1, p. 118, 2015

  180. [188]

    Surviving sepsis campaign: international guidelines for man- agement of sepsis and septic shock: 2016,

    A. Rhodes, L. E. Evans, W. Alhazzani, M. M. Levy, M. Antonelli, R. Ferrer, A. Kumar, J. E. Sevransky, C. L. Sprung, M. E. Nunnally et al. , “Surviving sepsis campaign: international guidelines for man- agement of sepsis and septic shock: 2016,” Intensive Care Medicine , vol. 4...

  181. [189]

    Acute respiratory distress syndrome,

    A. D. T. Force, V . Ranieri, G. Rubenfeld et al. , “Acute respiratory distress syndrome,” Jama, vol. 307, no. 23, pp. 2526–2533, 2012

  182. [190]

    Use of machine-learning approaches to predict clinical deterioration in critically ill patients: A systematic review,

    T. Kamio, T. Van, and K. Masamune, “Use of machine-learning approaches to predict clinical deterioration in critically ill patients: A systematic review,” International Journal of Medical Research and Health Sciences, vol. 6, no. 6, pp. 1–7, 2017

  183. [191]

    Machine learning in critical care: state-of-the-art and a sepsis case study,

    A. Vellido, V . Ribas, C. Morales, A. R. Sanmart ´ın, and J. C. R. Rodr´ıguez, “Machine learning in critical care: state-of-the-art and a sepsis case study,” Biomedical engineering online , vol. 17, no. 1, p. 135, 2018

  184. [192]

    A markov decision process to suggest optimal treatment of severe infections in intensive care,

    M. Komorowski, A. Gordon, L. Celi, and A. Faisal, “A markov decision process to suggest optimal treatment of severe infections in intensive care,” in Neural Information Processing Systems Workshop on Machine Learning for Health , 2016

  185. [193]

    The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care,

    M. Komorowski, L. A. Celi, O. Badawi, A. C. Gordon, and A. A. Faisal, “The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care,”Nature Medicine, vol. 24, no. 11, p. 1716, 2018

  186. [194]

    Deep reinforcement learning for sepsis treatment,

    A. Raghu, M. Komorowski, I. Ahmed, L. Celi, P. Szolovits, and M. Ghassemi, “Deep reinforcement learning for sepsis treatment,” arXiv preprint arXiv:1711.09602 , 2017

  187. [195]

    Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach,

    A. Raghu, M. Komorowski, L. A. Celi, P. Szolovits, and M. Ghassemi, “Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach,” in Machine Learning for Healthcare Conference, 2017, pp. 147–163

  188. [196]

    Model-based reinforcement learning for sepsis treatment,

    A. Raghu, M. Komorowski, and S. Singh, “Model-based reinforcement learning for sepsis treatment,” arXiv preprint arXiv:1811.09602, 2018

  189. [197]

    Treatment recommendation in crit- ical care: A scalable and interpretable approach in partially observable health states,

    C. P. Utomo, X. Li, and W. Chen, “Treatment recommendation in crit- ical care: A scalable and interpretable approach in partially observable health states,” 2018

  190. [198]

    Improving sepsis treatment strategies by combining deep and kernel-based reinforcement learning,

    X. Peng, Y . Ding, D. Wihl, O. Gottesman, M. Komorowski, L.-w. H. Lehman, A. Ross, A. Faisal, and F. Doshi-Velez, “Improving sepsis treatment strategies by combining deep and kernel-based reinforcement learning,” arXiv preprint arXiv:1901.04670 , 2019

  191. [199]

    Learning to treat sepsis with multi-output gaussian process deep recurrent q-networks,

    J. Futoma, A. Lin, M. Sendak, A. Bedoya, M. Clement, C. O’Brien, and K. Heller, “Learning to treat sepsis with multi-output gaussian process deep recurrent q-networks,” 2018

  192. [200]

    Deep inverse reinforcement learning for sepsis treatment,

    C. Yu, G. Ren, and J. Liu, “Deep inverse reinforcement learning for sepsis treatment,” in 2019 IEEE ICHI , 2019, pp. 1–3

  193. [201]

    The actor search tree critic (astc) for off-policy pomdp learning in medical decision making,

    L. Li, M. Komorowski, and A. A. Faisal, “The actor search tree critic (astc) for off-policy pomdp learning in medical decision making,”arXiv preprint arXiv:1805.11548, 2018

  194. [202]

    Representation and reinforcement learning for personalized glycemic control in septic patients,

    W.-H. Weng, M. Gao, Z. He, S. Yan, and P. Szolovits, “Representation and reinforcement learning for personalized glycemic control in septic patients,” arXiv preprint arXiv:1712.00654 , 2017

  195. [203]

    Precision medicine as a control problem: Using simulation and deep reinforcement learning to discover adap- tive, personalized multi-cytokine therapy for sepsis,

    B. K. Petersen, J. Yang, W. S. Grathwohl, C. Cockrell, C. Santiago, G. An, and D. M. Faissol, “Precision medicine as a control problem: Using simulation and deep reinforcement learning to discover adap- tive, personalized multi-cytokine therapy for sepsis,” arXiv preprint arXi...

  196. [204]

    Intelligent control of closed-loop sedation in simulated icu patients

    B. L. Moore, E. D. Sinzinger, T. M. Quasny, and L. D. Pyeatt, “Intelligent control of closed-loop sedation in simulated icu patients.” in FLAIRS Conference, 2004, pp. 109–114

  197. [205]

    Sedation of simulated icu patients using reinforcement learning based control,

    E. D. Sinzinger and B. Moore, “Sedation of simulated icu patients using reinforcement learning based control,” International Journal on Artificial Intelligence Tools, vol. 14, no. 01n02, pp. 137–156, 2005

  198. [206]

    Reinforcement learning: a novel method for optimal control of propofol-induced hypnosis,

    B. L. Moore, A. G. Doufas, and L. D. Pyeatt, “Reinforcement learning: a novel method for optimal control of propofol-induced hypnosis,” Anesthesia & Analgesia , vol. 112, no. 2, pp. 360–367, 2011

  199. [207]

    Reinforcement learning versus proportional–integral–derivative control of hypnosis in a simulated intraoperative patient,

    B. L. Moore, T. M. Quasny, and A. G. Doufas, “Reinforcement learning versus proportional–integral–derivative control of hypnosis in a simulated intraoperative patient,” Anesthesia & Analgesia , vol. 112, no. 2, pp. 350–359, 2011

  200. [208]

    Reinforcement learning for closed-loop propofol anesthesia: a study in human volunteers,

    B. L. Moore, L. D. Pyeatt, V . Kulkarni, P. Panousis, K. Padrez, and A. G. Doufas, “Reinforcement learning for closed-loop propofol anesthesia: a study in human volunteers,” The Journal of Machine Learning Research, vol. 15, no. 1, pp. 655–696, 2014

  201. [209]

    Reinforcement learning for closed-loop propofol anesthesia: A human volunteer study

    B. L. Moore, P. Panousis, V . Kulkarni, L. D. Pyeatt, and A. G. Doufas, “Reinforcement learning for closed-loop propofol anesthesia: A human volunteer study.” in IAAI, 2010

  202. [210]

    Multivariable anesthesia control using reinforcement learning,

    N. Sadati, A. Aflaki, and M. Jahed, “Multivariable anesthesia control using reinforcement learning,” in IEEE SMC’06, vol. 6. IEEE, 2006, pp. 4563–4568

  203. [211]

    An adaptive neural network filter for improved patient state estimation in closed- loop anesthesia control,

    E. C. Borera, B. L. Moore, A. G. Doufas, and L. D. Pyeatt, “An adaptive neural network filter for improved patient state estimation in closed- loop anesthesia control,” in IEEE ICTAI’11. IEEE, 2011, pp. 41–46

  204. [212]

    Towards efficient, personalized anesthesia using continuous reinforcement learning for propofol infusion control,

    C. Lowery and A. A. Faisal, “Towards efficient, personalized anesthesia using continuous reinforcement learning for propofol infusion control,” in IEEE/EMBS NER’13. IEEE, 2013, pp. 1414–1417

  205. [213]

    Closed-loop control of anesthesia and mean arterial pressure using reinforcement learning,

    R. Padmanabhan, N. Meskin, and W. M. Haddad, “Closed-loop control of anesthesia and mean arterial pressure using reinforcement learning,” Biomedical Signal Processing and Control , vol. 22, pp. 54–64, 2015

  206. [214]

    Learning from an expert

    P. Humbert, J. Audiffren, C. Dubost, and L. Oudre, “Learning from an expert.”

  207. [215]

    Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach,

    S. Nemati, M. M. Ghassemi, and G. D. Clifford, “Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach,” in IEEE 38th Annual International Conference of the Engineering in Medicine and Biology Society . IEEE, 2016, pp. 2978–2981

  208. [216]

    A deep deter- ministic policy gradient approach to medication dosing and surveillance in the icu,

    R. Lin, M. D. Stanley, M. M. Ghassemi, and S. Nemati, “A deep deter- ministic policy gradient approach to medication dosing and surveillance in the icu,” in IEEE EMBC’18. IEEE, 2018, pp. 4927–4931

  209. [217]

    Supervised reinforcement learning with recurrent neural network for dynamic treatment recom- mendation,

    L. Wang, W. Zhang, X. He, and H. Zha, “Supervised reinforcement learning with recurrent neural network for dynamic treatment recom- mendation,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . ACM, 2018, pp. 2447–2456

  210. [218]

    A reinforcement learning approach to weaning of mechanical ventilation in intensive care units,

    N. Prasad, L.-F. Cheng, C. Chivers, M. Draugelis, and B. E. Engel- hardt, “A reinforcement learning approach to weaning of mechanical ventilation in intensive care units,” arXiv preprint arXiv:1704.06300 , 2017

  211. [219]

    Inverse reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units,

    C. Yu, J. Liu, and H. Zhao, “Inverse reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units,” BMC medical informatics and decision making , vol. 19, no. 2, p. 57, 2019

  212. [220]

    Supervised-actor-critic reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units,

    C. Yu, G. Ren, and Y . Dong, “Supervised-actor-critic reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units,” BMC medical informatics and decision making , 2020

  213. [221]

    Towards high confidence off- policy reinforcement learning for clinical applications

    A. Jagannatha, P. Thomas, and H. Yu, “Towards high confidence off- policy reinforcement learning for clinical applications.”

  214. [222]

    An optimal policy for patient laboratory tests in intensive care units,

    L.-F. Cheng, N. Prasad, and B. E. Engelhardt, “An optimal policy for patient laboratory tests in intensive care units,” arXiv preprint arXiv:1808.04679, 2018

  215. [223]

    Dynamic measurement scheduling for adverse event forecasting using deep rl,

    C.-H. Chang, M. Mai, and A. Goldenberg, “Dynamic measurement scheduling for adverse event forecasting using deep rl,” arXiv preprint arXiv:1812.00268, 2018

  216. [224]

    Tools for the precision medicine era: How to develop highly personalized treatment recommendations from cohort and registry data using q-learning,

    E. F. Krakow, M. Hemmer, T. Wang, B. Logan, M. Arora, S. Spellman, D. Couriel, A. Alousi, J. Pidala, M. Last et al. , “Tools for the precision medicine era: How to develop highly personalized treatment recommendations from cohort and registry data using q-learning,” American j...

  217. [225]

    Deep reinforcement learning for dynamic treatment regimes on medical registry data,

    Y . Liu, B. Logan, N. Liu, Z. Xu, J. Tang, and Y . Wang, “Deep reinforcement learning for dynamic treatment regimes on medical registry data,” in IEEE ICHI’17. IEEE, 2017, pp. 380–385

  218. [226]

    Sepsis: pathophysiology and clinical management,

    J. E. Gotts and M. A. Matthay, “Sepsis: pathophysiology and clinical management,” Bmj, vol. 353, p. i1585, 2016

  219. [227]

    Mimic-iii, a freely accessible critical care database,

    A. E. Johnson, T. J. Pollard, L. Shen, H. L. Li-wei, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark, “Mimic-iii, a freely accessible critical care database,” Scientific Data, vol. 3, p. 160035, 2016

  220. [228]

    Individualized sepsis treatment using reinforcement learn- ing,

    S. Saria, “Individualized sepsis treatment using reinforcement learn- ing,” Nature medicine, vol. 24, no. 11, p. 1641, 2018

  221. [229]

    Deep reinforcement learning with double q-learning

    H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning.” in AAAI, vol. 2. Phoenix, AZ, 2016, p. 5

  222. [230]

    Dueling network architectures for deep reinforcement learning,

    Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in International Conference on Machine Learning , 2016, pp. 1995–2003

  223. [231]

    Prioritized experience replay,

    T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” arXiv preprint arXiv:1511.05952 , 2015

  224. [232]

    Prox- imal policy optimization algorithms,

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Prox- imal policy optimization algorithms,”arXiv preprint arXiv:1707.06347, 2017

  225. [233]

    Continuous control with deep reinforce- ment learning,

    T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y . Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforce- ment learning,” arXiv preprint arXiv:1509.02971 , 2015

  226. [234]

    Policy invariance under reward transformations: Theory and application to reward shaping,

    A. Y . Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in ICML, vol. 99, 1999, pp. 278–287

  227. [235]

    Clinical decision support and closed-loop control for intensive care unit sedation,

    W. M. Haddad, J. M. Bailey, B. Gholami, and A. R. Tannenbaum, “Clinical decision support and closed-loop control for intensive care unit sedation,” Asian Journal of Control, vol. 20, no. 5, pp. 1343–1350, 2012

  228. [236]

    A data-driven approach to optimized medication dosing: a focus on heparin,

    M. M. Ghassemi, S. E. Richter, I. M. Eche, T. W. Chen, J. Danziger, and L. A. Celi, “A data-driven approach to optimized medication dosing: a focus on heparin,” Intensive Care Medicine , vol. 40, no. 9, pp. 1332– 1339, 2014

  229. [237]

    The intensive care medicine research agenda for airways, invasive and noninvasive mechanical ventilation,

    S. Jaber, G. Bellani, L. Blanch, A. Demoule, A. Esteban, L. Gattinoni, C. Gu ´erin, N. Hill, J. G. Laffey, S. M. Maggiore et al., “The intensive care medicine research agenda for airways, invasive and noninvasive mechanical ventilation,” Intensive Care Medicine , vol. 43, no. ...

  230. [238]

    Focus on ventilation and airway management in the icu,

    A. De Jong, G. Citerio, and S. Jaber, “Focus on ventilation and airway management in the icu,” Intensive Care Medicine , vol. 43, no. 12, pp. 1912–1915, 2017

  231. [239]

    National Academies of Sciences, Medicine et al

    E. National Academies of Sciences, Medicine et al. , Improving diag- nosis in health care . National Academies Press, 2016

  232. [240]

    A review on use of machine learning techniques in diagnostic health-care,

    S. K. Rai and K. Sowmya, “A review on use of machine learning techniques in diagnostic health-care,” Artificial Intelligent Systems and Machine Learning, vol. 10, no. 4, pp. 102–107, 2018

  233. [241]

    Survey of machine learning algorithms for disease diagnostic,

    M. Fatima and M. Pasha, “Survey of machine learning algorithms for disease diagnostic,” Journal of Intelligent Learning Systems and Applications, vol. 9, no. 01, p. 1, 2017

  234. [242]

    Disease diagnosis in smart healthcare: Innovation, technologies and applications,

    K. T. Chui, W. Alhalabi, S. S. H. Pang, P. O. d. Pablos, R. W. Liu, and M. Zhao, “Disease diagnosis in smart healthcare: Innovation, technologies and applications,” Sustainability, vol. 9, no. 12, p. 2309, 2017

  235. [243]

    Learning to diagnose with lstm recurrent neural networks,

    Z. C. Lipton, D. C. Kale, C. Elkan, and R. Wetzel, “Learning to diagnose with lstm recurrent neural networks,” arXiv preprint arXiv:1511.03677, 2015

  236. [244]

    Retain: An interpretable predictive model for healthcare using reverse time attention mechanism,

    E. Choi, M. T. Bahadori, J. Sun, J. Kulas, A. Schuetz, and W. Stewart, “Retain: An interpretable predictive model for healthcare using reverse time attention mechanism,” in Advances in Neural Information Pro- cessing Systems, 2016, pp. 3504–3512

  237. [245]

    Medical question answering for clinical decision support,

    T. R. Goodwin and S. M. Harabagiu, “Medical question answering for clinical decision support,” in Proceedings of the 25th ACM Inter- national on Conference on Information and Knowledge Management . ACM, 2016, pp. 297–306

  238. [246]

    Diagnostic inferencing via improving clinical concept extraction with deep reinforcement learning: A preliminary study,

    Y . Ling, S. A. Hasan, V . Datla, A. Qadir, K. Lee, J. Liu, and O. Farri, “Diagnostic inferencing via improving clinical concept extraction with deep reinforcement learning: A preliminary study,” in Machine Learn- ing for Healthcare Conference , 2017, pp. 271–285

  239. [247]

    Reinforcement learning in computer vision,

    A. Bernstein and E. Burnaev, “Reinforcement learning in computer vision,” in CMV’17, vol. 10696. International Society for Optics and Photonics, 2018, p. 106961S

  240. [248]

    A reinforcement learning framework for parameter control in computer vision applications,

    G. W. Taylor, “A reinforcement learning framework for parameter control in computer vision applications,” in Computer and Robot Vision, 2004. Proceedings. First Canadian Conference on . IEEE, 2004, pp. 496–503

  241. [249]

    A reinforcement learning framework for medical image segmentation,

    F. Sahba, H. R. Tizhoosh, and M. M. Salama, “A reinforcement learning framework for medical image segmentation,” in IJCNN, vol. 6, 2006, pp. 511–517

  242. [250]

    Application of opposition-based reinforcement learning in image segmentation,

    ——, “Application of opposition-based reinforcement learning in image segmentation,” in 2007 IEEE Symposium on Computational Intelligence in Image and Signal Processing . IEEE, 2007, pp. 246– 251

  243. [251]

    Application of reinforcement learning for segmentation of transrectal ultrasound images,

    ——, “Application of reinforcement learning for segmentation of transrectal ultrasound images,” BMC Medical Imaging , vol. 8, no. 1, p. 8, 2008

  244. [252]

    Object segmentation in image sequences using reinforce- ment learning,

    F. Sahba, “Object segmentation in image sequences using reinforce- ment learning,” in CSCI’16. IEEE, 2016, pp. 1416–1417

  245. [253]

    Deep reinforcement learning for surgical gesture segmentation and classification,

    D. Liu and T. Jiang, “Deep reinforcement learning for surgical gesture segmentation and classification,” in International Conference on Med- ical Image Computing and Computer-Assisted Intervention . Springer, 2018, pp. 247–255

  246. [254]

    An artificial agent for anatomical landmark detection in medical images,

    F. C. Ghesu, B. Georgescu, T. Mansi, D. Neumann, J. Hornegger, and D. Comaniciu, “An artificial agent for anatomical landmark detection in medical images,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2016, pp. 229–237

  247. [255]

    Multi-scale deep reinforcement learning for real- time 3d-landmark detection in ct scans,

    F. C. Ghesu, B. Georgescu, Y . Zheng, S. Grbic, A. Maier, J. Hornegger, and D. Comaniciu, “Multi-scale deep reinforcement learning for real- time 3d-landmark detection in ct scans,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2017

  248. [256]

    Towards intelligent robust detection of anatomical structures in incomplete volumetric data,

    F. C. Ghesu, B. Georgescu, S. Grbic, A. Maier, J. Hornegger, and D. Comaniciu, “Towards intelligent robust detection of anatomical structures in incomplete volumetric data,” Medical Image Analysis , vol. 48, pp. 203–213, 2018

  249. [257]

    Nonlinear adaptively learned optimization for object localization in 3d medical images,

    M. Etcheverry, B. Georgescu, B. Odry, T. J. Re, S. Kaushik, B. Geiger, N. Mariappan, S. Grbic, and D. Comaniciu, “Nonlinear adaptively learned optimization for object localization in 3d medical images,” in Deep Learning in Medical Image Analysis and Multimodal Learning for Cli...

  250. [258]

    Evaluating reinforcement learning agents for anatomical landmark detection,

    A. Alansary, O. Oktay, Y . Li, L. Le Folgoc, B. Hou, G. Vaillant, B. Glocker, B. Kainz, and D. Rueckert, “Evaluating reinforcement learning agents for anatomical landmark detection,” 2018

  251. [259]

    Automatic view planning with multi-scale deep reinforcement learning agents,

    A. Alansary, L. L. Folgoc, G. Vaillant, O. Oktay, Y . Li, W. Bai, J. Passerat-Palmbach, R. Guerrero, K. Kamnitsas, B. Hou et al. , “Automatic view planning with multi-scale deep reinforcement learning agents,” arXiv preprint arXiv:1806.03228 , 2018

  252. [260]

    Partial policy-based reinforcement learning for anatomical landmark localization in 3d medical images,

    W. A. Al and I. D. Yun, “Partial policy-based reinforcement learning for anatomical landmark localization in 3d medical images,” arXiv preprint arXiv:1807.02908, 2018

  253. [261]

    An artificial agent for robust image registration

    R. Liao, S. Miao, P. de Tournemire, S. Grbic, A. Kamen, T. Mansi, and D. Comaniciu, “An artificial agent for robust image registration.” in AAAI, 2017, pp. 4168–4175

  254. [262]

    Multimodal image registration with deep context re- inforcement learning,

    K. Ma, J. Wang, V . Singh, B. Tamersoy, Y .-J. Chang, A. Wimmer, and T. Chen, “Multimodal image registration with deep context re- inforcement learning,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2017, pp. 240–248

  255. [263]

    Robust non-rigid registra- tion through agent-based action learning,

    J. Krebs, T. Mansi, H. Delingette, L. Zhang, F. C. Ghesu, S. Miao, A. K. Maier, N. Ayache, R. Liao, and A. Kamen, “Robust non-rigid registra- tion through agent-based action learning,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . ...

  256. [264]

    Deep reinforcement learning for active breast lesion detection from dce-mri,

    G. Maicas, G. Carneiro, A. P. Bradley, J. C. Nascimento, and I. Reid, “Deep reinforcement learning for active breast lesion detection from dce-mri,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2017, pp. 665–673

  257. [265]

    Deep reinforcement learning for vessel centerline tracing in multi-modality 3d volumes,

    P. Zhang, F. Wang, and Y . Zheng, “Deep reinforcement learning for vessel centerline tracing in multi-modality 3d volumes,” in Inter- national Conference on Medical Image Computing and Computer- Assisted Intervention. Springer, 2018, pp. 755–763

  258. [266]

    Application on reinforcement learning for diag- nosis based on medical image,

    S. M. B. Netto, V . R. C. Leite, A. C. Silva, A. C. de Paiva, and A. de Almeida Neto, “Application on reinforcement learning for diag- nosis based on medical image,” in Reinforcement Learning. InTech, 2008

  259. [267]

    Lead: a methodology for learning efficient approaches to medical diagnosis,

    S. J. Fakih and T. K. Das, “Lead: a methodology for learning efficient approaches to medical diagnosis,” IEEE Transactions on Information Technology in Biomedicine, vol. 10, no. 2, pp. 220–228, 2006

  260. [268]

    Overview of the trec 2016 clinical decision support track

    K. Roberts, M. S. Simpson, E. M. V oorhees, and W. R. Hersh, “Overview of the trec 2016 clinical decision support track.” in TREC, 2016

  261. [269]

    Learning to diagnose: Assimilating clinical narratives using deep reinforcement learning,

    Y . Ling, S. A. Hasan, V . Datla, A. Qadir, K. Lee, J. Liu, and O. Farri, “Learning to diagnose: Assimilating clinical narratives using deep reinforcement learning,” in Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Pape...

  262. [270]

    Breast cancer surveillance consortium: a national mammography screening and outcomes database

    R. Ballard-Barbash, S. H. Taplin, B. C. Yankaskas, V . L. Ernster, R. D. Rosenberg, P. A. Carney, W. E. Barlow, B. M. Geller, K. Kerlikowske, B. K. Edwards et al., “Breast cancer surveillance consortium: a national mammography screening and outcomes database.” American Journal...

  263. [271]

    An adaptive online learning framework for practical breast cancer diagnosis,

    T. Chu, J. Wang, and J. Chen, “An adaptive online learning framework for practical breast cancer diagnosis,” in Medical Imaging 2016: Computer-Aided Diagnosis, vol. 9785. International Society for Optics and Photonics, 2016, p. 978524

  264. [272]

    Inquire and di- agnose: Neural symptom checking ensemble using deep reinforcement learning,

    K.-F. Tang, H.-C. Kao, C.-N. Chou, and E. Y . Chang, “Inquire and di- agnose: Neural symptom checking ensemble using deep reinforcement learning,” in Proceedings of NIPS Workshop on Deep Reinforcement Learning, 2016

  265. [273]

    Context-aware symptom checking for disease diagnosis using hierarchical reinforcement learn- ing,

    H.-C. Kao, K.-F. Tang, and E. Y . Chang, “Context-aware symptom checking for disease diagnosis using hierarchical reinforcement learn- ing,” 2018

  266. [274]

    Artificial intelligence in xprize deepq tricorder,

    E. Y . Chang, M.-H. Wu, K.-F. T. Tang, H.-C. Kao, and C.-N. Chou, “Artificial intelligence in xprize deepq tricorder,” in Proceedings of the 2nd International Workshop on Multimedia for Personal Health and Health Care. ACM, 2017, pp. 11–18

  267. [275]

    Deepq: Advancing healthcare through artificial intel- ligence and virtual reality,

    E. Y . Chang, “Deepq: Advancing healthcare through artificial intel- ligence and virtual reality,” in Proceedings of the 2017 ACM on Multimedia Conference. ACM, 2017, pp. 1068–1068

  268. [276]

    Task-oriented dialogue system for automatic diagnosis,

    Z. Wei, Q. Liu, B. Peng, H. Tou, T. Chen, X. Huang, K.-F. Wong, and X. Dai, “Task-oriented dialogue system for automatic diagnosis,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , vol. 2, 2018, pp. 201–207

  269. [277]

    Improving mild cognitive impairment prediction via reinforcement learning and dialogue simulation,

    F. Tang, K. Lin, I. Uchendu, H. H. Dodge, and J. Zhou, “Improving mild cognitive impairment prediction via reinforcement learning and dialogue simulation,” arXiv preprint arXiv:1802.06428 , 2018

  270. [278]

    Approximate dynamic programming for capacity allocation in the service industry,

    H.-J. Schuetz and R. Kolisch, “Approximate dynamic programming for capacity allocation in the service industry,” European Journal of Operational Research, vol. 218, no. 1, pp. 239–250, 2012

  271. [279]

    Reinforcement learning based resource allocation in business process management,

    Z. Huang, W. M. van der Aalst, X. Lu, and H. Duan, “Reinforcement learning based resource allocation in business process management,” Data & Knowledge Engineering , vol. 70, no. 1, pp. 127–145, 2011

  272. [280]

    Clinic scheduling models with overbooking for patients with heterogeneous no-show probabilities,

    B. Zeng, A. Turkcan, J. Lin, and M. Lawley, “Clinic scheduling models with overbooking for patients with heterogeneous no-show probabilities,” Annals of Operations Research, vol. 178, no. 1, pp. 121– 144, 2010

  273. [281]

    Reinforcement learning for primary care e appointment scheduling,

    T. S. M. T. Gomes, “Reinforcement learning for primary care e appointment scheduling,” 2017

  274. [282]

    Asynchronous methods for deep reinforcement learning,

    V . Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in ICML, 2016, pp. 1928–1937

  275. [283]

    A function approximation method for model- based high-dimensional inverse reinforcement learning,

    K. Li and J. W. Burdick, “A function approximation method for model- based high-dimensional inverse reinforcement learning,” arXiv preprint arXiv:1708.07738, 2017

  276. [284]

    Multilateral surgical pattern cutting in 2d orthotropic gauze with deep reinforcement learning policies for tensioning,

    B. Thananjeyan, A. Garg, S. Krishnan, C. Chen, L. Miller, and K. Goldberg, “Multilateral surgical pattern cutting in 2d orthotropic gauze with deep reinforcement learning policies for tensioning,” in IEEE ICRA’17. IEEE, 2017, pp. 2371–2378

  277. [285]

    A new tensioning method using deep reinforcement learning for surgical pattern cutting,

    T. T. Nguyen, N. D. Nguyen, F. Bello, and S. Nahavandi, “A new tensioning method using deep reinforcement learning for surgical pattern cutting,” arXiv preprint arXiv:1901.03327 , 2019

  278. [286]

    Towards transferring skills to flexible surgical robots with programming by demonstration and reinforcement learning,

    J. Chen, H. Y . Lau, W. Xu, and H. Ren, “Towards transferring skills to flexible surgical robots with programming by demonstration and reinforcement learning,” in ICACI’16. IEEE, 2016, pp. 378–384

  279. [287]

    Path planning for automation of surgery robot based on probabilistic roadmap and reinforcement learning,

    D. Baek, M. Hwang, H. Kim, and D.-S. Kwon, “Path planning for automation of surgery robot based on probabilistic roadmap and reinforcement learning,” in 2018 15th International Conference on Ubiquitous Robots (UR) . IEEE, 2018, pp. 342–347

  280. [288]

    Inverse reinforcement learning via function approximation for clinical motion analysis,

    K. Li, M. Rath, and J. W. Burdick, “Inverse reinforcement learning via function approximation for clinical motion analysis,” in IEEE ICRA’18. IEEE, 2018, pp. 610–617

  281. [289]

    Training an actor-critic reinforcement learn- ing controller for arm movement using human-generated rewards,

    K. M. Jagodnik, P. S. Thomas, A. J. van den Bogert, M. S. Bran- icky, and R. F. Kirsch, “Training an actor-critic reinforcement learn- ing controller for arm movement using human-generated rewards,” IEEE Transactions on Neural Systems and Rehabilitation Engineering , vol. 25, ...

  282. [290]

    Medical qos provision based on reinforcement learning in ultrasound streaming over 3.5 g wireless systems,

    R. S. Istepanian, N. Y . Philip, and M. G. Martini, “Medical qos provision based on reinforcement learning in ultrasound streaming over 3.5 g wireless systems,” IEEE Journal on Selected areas in Communications, vol. 27, no. 4, 2009

  283. [291]

    Cross-layer ultrasound video streaming over mobile wimax and hsupa networks,

    A. Alinejad, N. Y . Philip, and R. S. Istepanian, “Cross-layer ultrasound video streaming over mobile wimax and hsupa networks,” IEEE transactions on Information Technology in Biomedicine, vol. 16, no. 1, pp. 31–39, 2012

  284. [292]

    Functional electrical stimulation after spinal cord injury: current use, therapeutic effects and future directions,

    K. Ragnarsson, “Functional electrical stimulation after spinal cord injury: current use, therapeutic effects and future directions,” Spinal cord, vol. 46, no. 4, p. 255, 2008

  285. [293]

    Creating a reinforcement learning controller for functional electrical stimulation of a human arm,

    P. S. Thomas, M. Branicky, A. Van Den Bogert, and K. Jagodnik, “Creating a reinforcement learning controller for functional electrical stimulation of a human arm,” in The Yale Workshop on Adaptive and Learning Systems, vol. 49326. NIH Public Access, 2008, p. 1

  286. [294]

    Application of the actor-critic architecture to functional electrical stimulation control of a human arm

    P. S. Thomas, A. J. van den Bogert, K. M. Jagodnik, and M. S. Branicky, “Application of the actor-critic architecture to functional electrical stimulation control of a human arm.” in IAAI, 2009

  287. [295]

    Can the pharmaceutical industry reduce attrition rates?

    I. Kola and J. Landis, “Can the pharmaceutical industry reduce attrition rates?” Nature reviews Drug discovery , vol. 3, no. 8, p. 711, 2004

  288. [296]

    Schneider, De novo molecular design

    G. Schneider, De novo molecular design . John Wiley & Sons, 2013

  289. [297]

    Molecular de-novo design through deep reinforcement learning,

    M. Olivecrona, T. Blaschke, O. Engkvist, and H. Chen, “Molecular de-novo design through deep reinforcement learning,” Journal of Cheminformatics, vol. 9, no. 1, p. 48, 2017

  290. [298]

    Accelerating drugs discovery with deep reinforcement learning: An early approach,

    A. Serrano, B. Imbern ´on, H. P ´erez-S´anchez, J. M. Cecilia, A. Bueno- Crespo, and J. L. Abell ´an, “Accelerating drugs discovery with deep reinforcement learning: An early approach,” in Proceedings of the 47th International Conference on Parallel Processing Companion . ACM,...

  291. [299]

    Exploring deep recurrent models with reinforcement learning for molecule design,

    D. Neil, M. Segler, L. Guasch, M. Ahmed, D. Plumbley, M. Sellwood, and N. Brown, “Exploring deep recurrent models with reinforcement learning for molecule design,” 2018

  292. [300]

    Deep reinforcement learning for de novo drug design,

    M. Popova, O. Isayev, and A. Tropsha, “Deep reinforcement learning for de novo drug design,” Science Advances, vol. 4, no. 7, p. eaap7885, 2018

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.