REVIEW 2 major objections 5 minor 2 cited by
Reinforcement Learning in Healthcare: A Survey
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Reinforcement learning works across healthcare, but its clinical value hinges on how states, rewards, and policies are formulated.
desk verdict Broad, well-organized RL-in-healthcare survey whose usefulness does not depend on the unsupported 'first comprehensive survey' claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central framework is the Markov decision process (MDP), defined as a 5-tuple $(S, A, P, R, \gamma)$, where an agent chooses actions to maximize discounted cumulative reward. The survey organizes RL techniques along two complementary directions: efficient techniques (experience replay and batch RL, model-based RL, transfer RL) and representational techniques (function approximation, deep RL, multi-objective RL, preference-based RL, inverse RL, factored MDPs, hierarchical RL, POMDPs).
What would settle it
A reader could falsify the survey's breadth claim by identifying a substantial published corpus of RL applications in healthcare that this survey omits, or by conducting a systematic literature search that yields a different distribution of application domains and open problems.
Extended reading notes
Core claim
The paper's central claim is that RL has been successfully applied across a broad range of healthcare domains, including dynamic treatment regimes for chronic diseases (cancer, diabetes, anemia, HIV, mental illness) and critical care (sepsis, anesthesia, mechanical ventilation), automated medical diagnosis from both structured and unstructured clinical data, and other domains such as health resource allocation, drug discovery, and health management. It argues that the RL framework—an agent learning optimal policies through trial-and-error interaction with an environment—maps naturally onto medical treatment as a sequential decision process. The authors assert that this survey is the first comprehensive survey of RL applications in healthcare.
Load-bearing premise
The claim that the survey is comprehensive and that its selected literature represents the field relies on an undocumented and potentially non-systematic search and inclusion process, rather than a transparent protocol.
Editorial extensions
If this is right
- If RL-based treatment policies prove reliable in real clinical settings, they could enable personalized, adaptive treatment regimens that account for individual patient heterogeneity.
- Successful integration of RL could reduce reliance on population-averaged treatment protocols by learning from accumulated patient data.
- Addressing reward formulation challenges could lead to treatment policies that balance competing objectives such as efficacy and toxicity more explicitly.
- Advances in off-policy evaluation and safe exploration would be needed before RL policies could be confidently deployed in clinical practice.
- Small-data and transfer learning techniques could extend RL to rare diseases or new patient cohorts where historical data is scarce.
Reading between the lines
- A testable extension would be a systematic benchmark that evaluates offline RL treatment policies against standardized clinical outcomes across multiple ICU databases, going beyond the selective studies the survey reports.
- The survey's emphasis on reward engineering suggests that progress may depend more on clinical expertise in specifying objectives than on advances in RL algorithms themselves.
- The paper's catalog of applications implies that interoperable standards for state representation and reward specification could accelerate cross-institution validation of RL-based clinical decision support.
- A consequence of the credit assignment discussion is that RL may be most easily validated in domains with short, well-defined decision horizons, such as medication dosing, before moving to longer-horizon interventions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a survey of reinforcement learning (RL) applications in healthcare. It begins with a review of RL foundations (MDPs, dynamic programming, Monte Carlo and temporal-difference methods, Q-learning, SARSA, policy search, actor-critic) and of key techniques (batch RL, model-based RL, transfer RL, representation learning for values/rewards/tasks, inverse RL, POMDPs). It then organizes the application literature into three broad areas: dynamic treatment regimes for chronic diseases and critical care, automated medical diagnosis from structured and unstructured data, and other healthcare domains such as resource allocation, process control, drug discovery, and health management. The final sections discuss challenges (state/action engineering, reward formulation, off-policy evaluation, model learning, exploration, credit assignment) and future directions (interpretability, prior knowledge, small data, ambient intelligence, and in-vivo validation). The conclusion states that the paper 'serves as the first comprehensive survey of RL applications in healthcare.'
Significance. If its coverage is accepted, the survey is a useful entry point for researchers new to RL in healthcare. Its strengths are the breadth of cited work, the clear application taxonomy, and the compact summary tables that map references to base methods, efficiency/representation techniques, data sources, and stated limitations. The treatment of off-policy evaluation, reward formulation, and the need for in-vivo validation is well aligned with the current state of the field. However, the paper's central contribution claim rests on 'first comprehensive survey,' and the manuscript provides no search protocol, inclusion/exclusion criteria, or comparison with prior surveys, so neither the priority nor the comprehensiveness can be independently verified. The paper also contains no limitation statement acknowledging this gap. These issues are load-bearing because comprehensive coverage is the stated basis of the survey's contribution, not merely a stylistic flourish.
major comments (2)
- [Section IX, with implications for Sections I and III-VI] The claim that the paper 'serves as the first comprehensive survey of RL applications in healthcare' is unsupported as written. The manuscript does not document a search protocol (databases, query terms, date range), inclusion/exclusion criteria, or any comparison with prior surveys, so a reader cannot check either 'first' or 'comprehensive.' Because this claim is the stated contribution of the paper, this is a load-bearing issue rather than a cosmetic one. I recommend either adding a methodology paragraph that makes the citation selection reproducible, or softening the claim to 'a survey' and explicitly stating the limitations of the coverage.
- [Section IX (conclusion) and the absence of a survey-limitations paragraph] The paper nowhere acknowledges the limitations of its own survey methodology. In particular, it does not state that the application map may be incomplete, that the selected references may be biased toward work visible to the authors, or that the catalog of challenges in Section VII is a synthesis of a non-systematically selected subset of the literature. Given that the paper explicitly builds its contribution on comprehensiveness, the absence of such a limitation statement should be corrected in the revision.
minor comments (5)
- [Section II-A, Eq. (7)] The epsilon-greedy policy in Eq. (7) is not a valid probability distribution: it assigns probability 1-epsilon to the greedy action and probability epsilon to every non-greedy action, which sums to more than 1 when the action set has more than two actions. The standard formulation is 1-epsilon+epsilon/|A| for the greedy action and epsilon/|A| for each non-greedy action.
- [Section IV-A1 and Table III] The acronym ODE is defined as 'Ordinary Difference Equations' in the text, but the cited models are systems of ordinary differential equations; the definition should read 'Ordinary Differential Equations.'
- [Section IV-A1, cancer chemotherapy paragraph] The text refers to 'support vector regression (SVG)' while Table III and the reference [104] describe support vector regression (SVR); the acronym should be corrected consistently to SVR.
- [Section II-A, Eq. (6)] The SARSA update in Eq. (6) is written as Q_t(s', pi(s')), but SARSA updates use the actually sampled next action a' drawn from the behavior policy; as written the notation conflates the behavior policy with the target policy and could confuse readers.
- [Section IV-B and Table IV] The text uses '3C (Compartimentalization, Corruption, and Complexity)'; the intended term appears to be 'Compartmentalization,' and the spelling should be corrected.
Circularity Check
No circularity: this survey's taxonomy and challenge synthesis are independent of its inputs, and its authors' self-citations are descriptive rather than load-bearing.
full rationale
This paper is a survey, not a derivation. It does not claim to derive a novel result from first principles or to predict an outcome from fitted parameters. Its central content is an enumeration and organization of existing RL applications in healthcare, supported by external studies with their own methods, data, and evaluations. The organizational claims in Sections III-VI and the challenge lists in Sections VII-VIII are self-contained syntheses of the cited literature rather than reductions to the survey's own inputs. The authors' self-citations, such as [149] on causal policy gradient for HIV, [200] on deep IRL for sepsis treatment, [219] on inverse RL for mechanical ventilation, and [220] on supervised actor-critic for ventilation, appear as examples among many other works and are not used to justify the survey's conclusions. The 'first comprehensive survey' claim in Section IX is a completeness and historical claim; the absence of a documented search protocol is a coverage or methodological limitation, not a circularity defect. No fitted parameter is renamed as a prediction, no equation is defined in terms of its own output, and no load-bearing argument reduces to a self-citation. Therefore, per the default expectation for surveys, the circularity score is 0.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Reinforcement Learning in Healthcare: A Survey." pith.science (2026). https://pith.science/paper/SOG4EKXG
@misc{pith2026190808796,
author = {Pith},
title = {Pith review of: Reinforcement Learning in Healthcare: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/SOG4EKXG}},
note = {Machine review of arXiv:1908.08796}
}
read the original abstract
As a subfield of machine learning, reinforcement learning (RL) aims at empowering one's capabilities in behavioural decision making by using interaction experience with the world and an evaluative feedback. Unlike traditional supervised learning methods that usually rely on one-shot, exhaustive and supervised reward signals, RL tackles with sequential decision making problems with sampled, evaluative and delayed feedback simultaneously. Such distinctive features make RL technique a suitable candidate for developing powerful solutions in a variety of healthcare domains, where diagnosing decisions or treatment regimes are usually characterized by a prolonged and sequential procedure. This survey discusses the broad applications of RL techniques in healthcare domains, in order to provide the research community with systematic understanding of theoretical foundations, enabling methods and techniques, existing challenges, and new insights of this emerging paradigm. By first briefly examining theoretical foundations and key techniques in RL research from efficient and representational directions, we then provide an overview of RL applications in healthcare domains ranging from dynamic treatment regimes in chronic diseases and critical care, automated medical diagnosis from both unstructured and structured clinical data, as well as many other control or scheduling domains that have infiltrated many aspects of a healthcare system. Finally, we summarize the challenges and open issues in current research, and point out some potential solutions and directions for future research.
Figures
Forward citations
Cited by 2 Pith papers
-
Counterfactual Shapley Credit Assignment
Counterfactual Shapley values, computed by simulated 'what-if' action replacements, redistribute RL rewards without changing the optimal policy and improve credit assignment in stochastic, sparse, delayed-reward tasks.
-
Reinforcement Learning for a Discrete-Time Linear-Quadratic Control Problem with an Application
The paper claims entropy regularization forces the optimal LQ feedback policy to be Gaussian and uses that to solve a mean-variance asset-liability problem, but the proof of the main theorem contains a correlation err...
Reference graph
Works this paper leans on
-
[1]
The coming of age of artificial intelligence in medicine,
V . L. Patel, E. H. Shortliffe, M. Stefanelli, P. Szolovits, M. R. Berthold, R. Bellazzi, and A. Abu-Hanna, “The coming of age of artificial intelligence in medicine,” Artificial Intelligence in Medicine , vol. 46, no. 1, pp. 5–17, 2009
2009
-
[2]
Artificial intelligence in medicine and cardiac imaging: harnessing big data and advanced computing to provide personalized medical diagnosis and treatment,
S. E. Dilsizian and E. L. Siegel, “Artificial intelligence in medicine and cardiac imaging: harnessing big data and advanced computing to provide personalized medical diagnosis and treatment,” Current Cardiology Reports, vol. 16, no. 1, p. 441, 2014
2014
-
[3]
Artificial intelligence in healthcare: past, present and future,
F. Jiang, Y . Jiang, H. Zhi, Y . Dong, H. Li, S. Ma, Y . Wang, Q. Dong, H. Shen, and Y . Wang, “Artificial intelligence in healthcare: past, present and future,” Stroke and Vascular Neurology, vol. 2, no. 4, pp. 230–243, 2017
2017
-
[4]
The practical implementation of artificial intelligence technologies in medicine,
J. He, S. L. Baxter, J. Xu, J. Xu, X. Zhou, and K. Zhang, “The practical implementation of artificial intelligence technologies in medicine,” Nature Medicine, vol. 25, no. 1, p. 30, 2019
2019
-
[5]
Machine learning and decision support in critical care,
A. E. Johnson, M. M. Ghassemi, S. Nemati, K. E. Niehaus, D. A. Clifton, and G. D. Clifford, “Machine learning and decision support in critical care,” Proceedings of the IEEE , vol. 104, no. 2, pp. 444–466, 2016
2016
-
[6]
Deep learning for health informatics,
D. Rav `ı, C. Wong, F. Deligianni, M. Berthelot, J. Andreu-Perez, B. Lo, and G.-Z. Yang, “Deep learning for health informatics,” IEEE Journal of Biomedical and Health Informatics , vol. 21, no. 1, pp. 4–21, 2017
2017
-
[7]
Opportunities and obstacles for deep learning in biology and medicine,
T. Ching, D. S. Himmelstein, B. K. Beaulieu-Jones, A. A. Kalinin, B. T. Do, G. P. Way, E. Ferrero, P.-M. Agapow, M. Zietz, M. M. Hoffman et al. , “Opportunities and obstacles for deep learning in biology and medicine,” bioRxiv, p. 142760, 2018
2018
-
[8]
Big data application in biomedical research and health care: a literature review,
J. Luo, M. Wu, D. Gopukumar, and Y . Zhao, “Big data application in biomedical research and health care: a literature review,” Biomedical Informatics Insights, vol. 8, pp. BII–S31 559, 2016
2016
Show all 300 references
-
[9]
A guide to deep learning in healthcare,
A. Esteva, A. Robicquet, B. Ramsundar, V . Kuleshov, M. DePristo, K. Chou, C. Cui, G. Corrado, S. Thrun, and J. Dean, “A guide to deep learning in healthcare,” Nature Medicine, vol. 25, no. 1, p. 24, 2019
2019
-
[10]
Human-level control through deep reinforcement learning,
V . Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, p. 529, 2015
2015
-
[11]
Reinforcement learning improves behaviour from evaluative feedback,
M. L. Littman, “Reinforcement learning improves behaviour from evaluative feedback,” Nature, vol. 521, no. 7553, p. 445, 2015
2015
-
[12]
Deep reinforcement learning,
Y . Li, “Deep reinforcement learning,”arXiv preprint arXiv:1810.06339, 2018
2018 arXiv
-
[13]
Applications of deep learning and reinforcement learning to biological data,
M. Mahmud, M. S. Kaiser, A. Hussain, and S. Vassanelli, “Applications of deep learning and reinforcement learning to biological data,” IEEE transactions on neural networks and learning systems , vol. 29, no. 6, pp. 2063–2079, 2018
2018
-
[14]
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction. MIT press, 2018
2018
-
[15]
Re- inforcement learning for control: Performance, stability, and deep approximators,
L. Bus ¸oniu, T. de Bruin, D. Toli ´c, J. Kober, and I. Palunko, “Re- inforcement learning for control: Performance, stability, and deep approximators,” Annual Reviews in Control , 2018
2018
-
[16]
Guidelines for reinforcement learning in healthcare
O. Gottesman, F. Johansson, M. Komorowski, A. Faisal, D. Sontag, F. Doshi-Velez, and L. A. Celi, “Guidelines for reinforcement learning in healthcare.” Nature medicine, vol. 25, no. 1, p. 16, 2019
2019
-
[17]
Bellman, Dynamic programming
R. Bellman, Dynamic programming. Courier Corporation, 2013
2013
-
[18]
Q-learning,
C. J. Watkins and P. Dayan, “Q-learning,” Machine Learning, vol. 8, no. 3-4, pp. 279–292, 1992
1992
-
[19]
G. A. Rummery and M. Niranjan, On-line Q-learning using connec- tionist systems. University of Cambridge, Department of Engineering Cambridge, England, 1994, vol. 37
1994
-
[20]
Policy search for motor primitives in robotics,
J. Kober and J. R. Peters, “Policy search for motor primitives in robotics,” inAdvances in Neural Information Processing Systems, 2009, pp. 849–856
2009
-
[21]
Natural actor-critic,
J. Peters and S. Schaal, “Natural actor-critic,” Neurocomputing, vol. 71, no. 7-9, pp. 1180–1190, 2008
2008
-
[22]
Bayesian reinforcement learning,
N. Vlassis, M. Ghavamzadeh, S. Mannor, and P. Poupart, “Bayesian reinforcement learning,” in Reinforcement Learning. Springer, 2012, pp. 359–386
2012
-
[23]
Bayesian reinforcement learning: A survey,
M. Ghavamzadeh, S. Mannor, J. Pineau, A. Tamar et al. , “Bayesian reinforcement learning: A survey,” Foundations and Trends R© in Ma- chine Learning, vol. 8, no. 5-6, pp. 359–483, 2015
2015
-
[24]
Reinforcement learning in finite mdps: Pac analysis,
A. L. Strehl, L. Li, and M. L. Littman, “Reinforcement learning in finite mdps: Pac analysis,” Journal of Machine Learning Research , vol. 10, no. Nov, pp. 2413–2444, 2009
2009
-
[25]
R-max-a general polynomial time algorithm for near-optimal reinforcement learning,
R. I. Brafman and M. Tennenholtz, “R-max-a general polynomial time algorithm for near-optimal reinforcement learning,”Journal of Machine Learning Research, vol. 3, no. Oct, pp. 213–231, 2002
2002
-
[26]
Intrinsic motivation and reinforcement learning,
A. G. Barto, “Intrinsic motivation and reinforcement learning,” in Intrinsically motivated learning in natural and artificial systems . Springer, 2013, pp. 17–47
2013
-
[27]
A comprehensive survey of multiagent reinforcement learning,
L. Busoniu, R. Babuska, and B. De Schutter, “A comprehensive survey of multiagent reinforcement learning,” IEEE Transactions on Systems, Man, And Cybernetics-Part C: Applications and Reviews, 38 (2), 2008 , 2008
2008
-
[28]
On the sample complexity of reinforcement learning,
S. M. Kakade et al. , “On the sample complexity of reinforcement learning,” Ph.D. dissertation, University of London London, England, 2003
2003
-
[29]
Sample complexity bounds of exploration,
L. Li, “Sample complexity bounds of exploration,” in Reinforcement Learning. Springer, 2012, pp. 175–204
2012
-
[30]
Busoniu, R
L. Busoniu, R. Babuska, B. De Schutter, and D. Ernst, Reinforcement learning and dynamic programming using function approximators . CRC press, 2010
2010
-
[31]
Reinforcement learning in continuous state and action spaces,
H. Van Hasselt, “Reinforcement learning in continuous state and action spaces,” in Reinforcement learning. Springer, 2012, pp. 207–251
2012
-
[32]
Safe exploration in markov decision processes,
T. M. Moldovan and P. Abbeel, “Safe exploration in markov decision processes,” in Proceedings of the 29th International Coference on International Conference on Machine Learning . Omnipress, 2012, pp. 1451–1458
2012
-
[33]
A comprehensive survey on safe rein- forcement learning,
J. Garcıa and F. Fern ´andez, “A comprehensive survey on safe rein- forcement learning,” Journal of Machine Learning Research , vol. 16, no. 1, pp. 1437–1480, 2015
2015
-
[34]
Robust markov decision processes,
W. Wiesemann, D. Kuhn, and B. Rustem, “Robust markov decision processes,” Mathematics of Operations Research , vol. 38, no. 1, pp. 153–183, 2013
2013
-
[35]
Distributionally robust markov decision pro- cesses,
H. Xu and S. Mannor, “Distributionally robust markov decision pro- cesses,” in Advances in Neural Information Processing Systems , 2010, pp. 2505–2513
2010
-
[36]
Interpretable policies for rein- forcement learning by genetic programming,
D. Hein, S. Udluft, and T. A. Runkler, “Interpretable policies for rein- forcement learning by genetic programming,”Engineering Applications of Artificial Intelligence , vol. 76, pp. 158–169, 2018
2018
-
[37]
Verifiable reinforcement learning via policy extraction,
O. Bastani, Y . Pu, and A. Solar-Lezama, “Verifiable reinforcement learning via policy extraction,” in Advances in Neural Information Processing Systems, 2018, pp. 2499–2509
2018
-
[38]
Reinforcement learning,
M. Wiering and M. Van Otterlo, “Reinforcement learning,” Adaptation, learning, and optimization , vol. 12, 2012
2012
-
[39]
Batch reinforcement learning,
S. Lange, T. Gabel, and M. Riedmiller, “Batch reinforcement learning,” in Reinforcement learning. Springer, 2012, pp. 45–73
2012
-
[40]
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method,
M. Riedmiller, “Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method,” in European Confer- ence on Machine Learning . Springer, 2005, pp. 317–328
2005
-
[41]
Tree-based batch mode rein- forcement learning,
D. Ernst, P. Geurts, and L. Wehenkel, “Tree-based batch mode rein- forcement learning,” Journal of Machine Learning Research , vol. 6, no. Apr, pp. 503–556, 2005
2005
-
[42]
Least-squares policy iteration,
M. G. Lagoudakis and R. Parr, “Least-squares policy iteration,” Journal of Machine Learning Research , vol. 4, no. Dec, pp. 1107–1149, 2003
2003
-
[43]
Learning and using models,
T. Hester and P. Stone, “Learning and using models,” in Reinforcement learning. Springer, 2012, pp. 111–141
2012
-
[44]
A survey of monte carlo tree search methods,
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton, “A survey of monte carlo tree search methods,” IEEE Transactions on Computational Intelligence and AI in games , vol. 4, no. 1, pp. 1–43, 2012
2012
-
[45]
Transfer in reinforcement learning: a framework and a survey,
A. Lazaric, “Transfer in reinforcement learning: a framework and a survey,” in Reinforcement Learning. Springer, 2012, pp. 143–173
2012
-
[46]
Transfer learning for reinforcement learning domains: A survey,
M. E. Taylor and P. Stone, “Transfer learning for reinforcement learning domains: A survey,” Journal of Machine Learning Research , vol. 10, no. Jul, pp. 1633–1685, 2009
2009
-
[47]
Mastering the game of go with deep neural networks and tree search,
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V . Panneershelvam, M. Lanctot et al., “Mastering the game of go with deep neural networks and tree search,” Nature, vol. 529, no. 7587, p. 484, 2016
2016
-
[48]
A survey of deep neural network architectures and their applications,
W. Liu, Z. Wang, X. Liu, N. Zeng, Y . Liu, and F. E. Alsaadi, “A survey of deep neural network architectures and their applications,” Neurocomputing, vol. 234, pp. 11–26, 2017
2017
-
[49]
Efficient processing of deep neural networks: A tutorial and survey,
V . Sze, Y .-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient processing of deep neural networks: A tutorial and survey,” Proceedings of the IEEE, vol. 105, no. 12, pp. 2295–2329, 2017
2017
-
[50]
Multiobjective reinforcement learning: A comprehensive overview,
C. Liu, X. Xu, and D. Hu, “Multiobjective reinforcement learning: A comprehensive overview,” IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 45, no. 3, pp. 385–398, 2015
2015
-
[51]
A survey of preference-based reinforcement learning methods,
C. Wirth, R. Akrour, G. Neumann, and J. F ¨urnkranz, “A survey of preference-based reinforcement learning methods,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 4945–4990, 2017
2017
-
[52]
Preference- based reinforcement learning: a formal framework and a policy iteration algorithm,
J. F ¨urnkranz, E. H ¨ullermeier, W. Cheng, and S.-H. Park, “Preference- based reinforcement learning: a formal framework and a policy iteration algorithm,” Machine Learning, vol. 89, no. 1-2, pp. 123–156, 2012
2012
-
[53]
Algorithms for inverse reinforcement learning
A. Y . Ng, S. J. Russell et al. , “Algorithms for inverse reinforcement learning.” in ICML, 2000, pp. 663–670
2000
-
[54]
A survey of inverse reinforcement learning techniques,
S. Zhifei and E. Meng Joo, “A survey of inverse reinforcement learning techniques,” International Journal of Intelligent Computing and Cybernetics, vol. 5, no. 3, pp. 293–311, 2012
2012
-
[55]
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning.” in AAAI, vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438
2008
-
[56]
Apprenticeship learning via inverse rein- forcement learning,
P. Abbeel and A. Y . Ng, “Apprenticeship learning via inverse rein- forcement learning,” in Proceedings of the twenty-first international conference on Machine learning . ACM, 2004, p. 1
2004
-
[57]
Nonlinear inverse reinforcement learning with gaussian processes,
S. Levine, Z. Popovic, and V . Koltun, “Nonlinear inverse reinforcement learning with gaussian processes,” in Advances in Neural Information Processing Systems, 2011, pp. 19–27
2011
-
[58]
Bayesian inverse reinforcement learn- ing,
D. Ramachandran and E. Amir, “Bayesian inverse reinforcement learn- ing,” Urbana, vol. 51, no. 61801, pp. 1–4, 2007
2007
-
[59]
Efficient solution algorithms for factored mdps,
C. Guestrin, D. Koller, R. Parr, and S. Venkataraman, “Efficient solution algorithms for factored mdps,”Journal of Artificial Intelligence Research, vol. 19, pp. 399–468, 2003
2003
-
[60]
Efficient reinforcement learning in factored mdps,
M. Kearns and D. Koller, “Efficient reinforcement learning in factored mdps,” in IJCAI, vol. 16, 1999, pp. 740–747
1999
-
[61]
Algorithm-directed exploration for model-based reinforcement learning in factored mdps,
C. Guestrin, R. Patrascu, and D. Schuurmans, “Algorithm-directed exploration for model-based reinforcement learning in factored mdps,” in ICML, 2002, pp. 235–242
2002
-
[62]
Near-optimal reinforcement learning in factored mdps,
I. Osband and B. Van Roy, “Near-optimal reinforcement learning in factored mdps,” inAdvances in Neural Information Processing Systems, 2014, pp. 604–612
2014
-
[63]
Efficient structure learning in factored-state mdps,
A. L. Strehl, C. Diuk, and M. L. Littman, “Efficient structure learning in factored-state mdps,” in AAAI, vol. 7, 2007, pp. 645–650
2007
-
[64]
Recent advances in hierarchical reinforcement learning,
A. G. Barto and S. Mahadevan, “Recent advances in hierarchical reinforcement learning,” Discrete Event Dynamic Systems , vol. 13, no. 1-2, pp. 41–77, 2003
2003
-
[65]
Hierarchical approaches,
B. Hengst, “Hierarchical approaches,” in Reinforcement learning . Springer, 2012, pp. 293–323
2012
-
[66]
Solving relational and first-order logical markov decision processes: A survey,
M. van Otterlo, “Solving relational and first-order logical markov decision processes: A survey,” in Reinforcement Learning. Springer, 2012, pp. 253–292
2012
-
[67]
Reinforcement learning algorithm for partially observable markov decision problems,
T. Jaakkola, S. P. Singh, and M. I. Jordan, “Reinforcement learning algorithm for partially observable markov decision problems,” in Ad- vances in Neural Information Processing Systems , 1995, pp. 345–352
1995
-
[68]
A computer program for digitalis dosage regimens,
R. W. Jelliffe, J. Buell, R. Kalaba, R. Sridhar, and R. Rockwell, “A computer program for digitalis dosage regimens,” Mathematical Biosciences, vol. 9, pp. 179–193, 1970
1970
-
[69]
R. E. Bellman, Mathematical methods in medicine . World Scientific Publishing Co., Inc., 1983
1983
-
[70]
Comparison of some control strategies for three-compartment pk/pd models,
C. Hu, W. S. Lovejoy, and S. L. Shafer, “Comparison of some control strategies for three-compartment pk/pd models,” Journal of Pharmacokinetics and Biopharmaceutics , vol. 22, no. 6, pp. 525–550, 1994
1994
-
[71]
Modeling medical treatment using markov decision processes,
A. J. Schaefer, M. D. Bailey, S. M. Shechter, and M. S. Roberts, “Modeling medical treatment using markov decision processes,” in Operations Research and Health Care . Springer, 2005, pp. 593–612
2005
-
[72]
Dynamic treatment regimes,
B. Chakraborty and S. A. Murphy, “Dynamic treatment regimes,” Annual Review of Statistics and Its Application , vol. 1, pp. 447–464, 2014
2014
-
[73]
Dynamic treatment regimes: Technical challenges and applications,
E. B. Laber, D. J. Lizotte, M. Qian, W. E. Pelham, and S. A. Murphy, “Dynamic treatment regimes: Technical challenges and applications,” Electronic Journal of Statistics , vol. 8, no. 1, p. 1225, 2014
2014
-
[74]
Estimation of survival distributions of treatment policies in two-stage randomization designs in clinical trials,
J. K. Lunceford, M. Davidian, and A. A. Tsiatis, “Estimation of survival distributions of treatment policies in two-stage randomization designs in clinical trials,” Biometrics, vol. 58, no. 1, pp. 48–57, 2002
2002
-
[75]
Adaptive interventions in child and adolescent mental health,
D. Almirall and A. Chronis-Tuscano, “Adaptive interventions in child and adolescent mental health,” Journal of Clinical Child & Adolescent Psychology, vol. 45, no. 4, pp. 383–395, 2016
2016
-
[76]
Adaptive treatment strategies in chronic disease,
P. W. Lavori and R. Dawson, “Adaptive treatment strategies in chronic disease,” Annu. Rev. Med., vol. 59, pp. 443–453, 2008
2008
-
[77]
Chakraborty and E
B. Chakraborty and E. E. M. Moodie, Statistical Reinforcement Learn- ing. Springer New York, 2013
2013
-
[78]
An experimental design for the development of adaptive treatment strategies,
S. A. Murphy, “An experimental design for the development of adaptive treatment strategies,” Statistics in Medicine , vol. 24, no. 10, pp. 1455– 1481, 2005
2005
-
[79]
Developing adaptive treatment strategies in substance abuse research,
S. A. Murphy, K. G. Lynch, D. Oslin, J. R. McKay, and T. TenHave, “Developing adaptive treatment strategies in substance abuse research,” Drug & Alcohol Dependence , vol. 88, pp. S24–S30, 2007
2007
-
[80]
W. H. Organization, Preventing chronic diseases: a vital investment . World Health Organization, 2005
2005
-
[81]
Chakraborty and E
B. Chakraborty and E. Moodie, Statistical methods for dynamic treat- ment regimes. Springer, 2013
2013
-
[82]
Improving chronic illness care: translating evidence into action,
E. H. Wagner, B. T. Austin, C. Davis, M. Hindmarsh, J. Schaefer, and A. Bonomi, “Improving chronic illness care: translating evidence into action,” Health Affairs, vol. 20, no. 6, pp. 64–78, 2001
2001
-
[83]
Reinforcement learning design for cancer clinical trials,
Y . Zhao, M. R. Kosorok, and D. Zeng, “Reinforcement learning design for cancer clinical trials,” Statistics in Medicine , vol. 28, no. 26, pp. 3294–3315, 2009
2009
-
[84]
Reinforcement learning based control of tumor growth with chemotherapy,
A. Hassani et al. , “Reinforcement learning based control of tumor growth with chemotherapy,” in 2010 International Conference on System Science and Engineering (ICSSE) . IEEE, 2010, pp. 185–189
2010
-
[85]
Drug scheduling of cancer chemotherapy based on natural actor-critic approach,
I. Ahn and J. Park, “Drug scheduling of cancer chemotherapy based on natural actor-critic approach,” BioSystems, vol. 106, no. 2-3, pp. 121–129, 2011
2011
-
[86]
Using reinforcement learning to personalize dosing strategies in a simulated cancer trial with high dimensional data,
K. Humphrey, “Using reinforcement learning to personalize dosing strategies in a simulated cancer trial with high dimensional data,” 2017
2017
-
[87]
Reinforcement learning-based control of drug dosing for cancer chemotherapy treat- ment,
R. Padmanabhan, N. Meskin, and W. M. Haddad, “Reinforcement learning-based control of drug dosing for cancer chemotherapy treat- ment,” Mathematical biosciences, vol. 293, pp. 11–20, 2017
2017
-
[88]
Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer,
Y . Zhao, D. Zeng, M. A. Socinski, and M. R. Kosorok, “Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer,” Biometrics, vol. 67, no. 4, pp. 1422–1433, 2011
2011
-
[89]
Preference- based policy iteration: Leveraging preference learning for reinforce- ment learning,
W. Cheng, J. F ¨urnkranz, E. H ¨ullermeier, and S.-H. Park, “Preference- based policy iteration: Leveraging preference learning for reinforce- ment learning,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2011, pp. 312– 327
2011
-
[90]
April: Active preference learning-based reinforcement learning,
R. Akrour, M. Schoenauer, and M. Sebag, “April: Active preference learning-based reinforcement learning,” in Joint European Confer- ence on Machine Learning and Knowledge Discovery in Databases . Springer, 2012, pp. 116–131
2012
-
[91]
Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm,
R. Busa-Fekete, B. Sz ¨or´enyi, P. Weng, W. Cheng, and E. H ¨ullermeier, “Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm,” Machine Learning, vol. 97, no. 3, pp. 327–351, 2014
2014
-
[92]
Reinforcement learning in models of adaptive medical treatment strategies,
R. Vincent, “Reinforcement learning in models of adaptive medical treatment strategies,” Ph.D. dissertation, McGill University Libraries, 2014
2014
-
[93]
Deep reinforcement learning for automated radiation adaptation in lung cancer,
H. H. Tseng, Y . Luo, S. Cui, J. T. Chien, R. K. Ten Haken, and I. E. Naqa, “Deep reinforcement learning for automated radiation adaptation in lung cancer,”Medical Physics, vol. 44, no. 12, pp. 6690–6705, 2017
2017
-
[94]
Simulation-based optimization of radiotherapy: Agent-based modeling and reinforcement learning,
A. Jalalimanesh, H. S. Haghighi, A. Ahmadi, and M. Soltani, “Simulation-based optimization of radiotherapy: Agent-based modeling and reinforcement learning,” Mathematics and Computers in Simula- tion, vol. 133, pp. 235–248, 2017
2017
-
[95]
Multi-objective optimization of radiotherapy: distributed q-learning and agent-based simulation,
A. Jalalimanesh, H. S. Haghighi, A. Ahmadi, H. Hejazian, and M. Soltani, “Multi-objective optimization of radiotherapy: distributed q-learning and agent-based simulation,” Journal of Experimental & Theoretical Artificial Intelligence , pp. 1–16, 2017
2017
-
[96]
Q-learning with censored data,
Y . Goldberg and M. R. Kosorok, “Q-learning with censored data,” Annals of Statistics , vol. 40, no. 1, p. 529, 2012
2012
-
[97]
Personalized medical treatments using novel reinforce- ment learning algorithms,
Y . M. Soliman, “Personalized medical treatments using novel reinforce- ment learning algorithms,” arXiv preprint arXiv:1406.3922 , 2014
2014 arXiv
-
[98]
Reinforcement learning with action-derived rewards for chemotherapy and clinical trial dosing regimen selection,
G. Yauney and P. Shah, “Reinforcement learning with action-derived rewards for chemotherapy and clinical trial dosing regimen selection,” in Machine Learning for Healthcare Conference , 2018, pp. 161–226
2018
-
[99]
World cancer report 2014,
B. Stewart, C. P. Wild et al., “World cancer report 2014,”Health, 2017
2014
-
[100]
Interactions between the immune system and cancer: a brief review of non-spatial mathematical models,
R. Eftimie, J. L. Bramson, and D. J. Earn, “Interactions between the immune system and cancer: a brief review of non-spatial mathematical models,” Bulletin of Mathematical Biology , vol. 73, no. 1, pp. 2–32, 2011
2011
-
[101]
A survey of optimiza- tion models on cancer chemotherapy treatment planning,
J. Shi, O. Alagoz, F. S. Erenay, and Q. Su, “A survey of optimiza- tion models on cancer chemotherapy treatment planning,” Annals of Operations Research, vol. 221, no. 1, pp. 331–356, 2014
2014
-
[102]
Cancer evolution: mathematical models and computational inference,
N. Beerenwinkel, R. F. Schwarz, M. Gerstung, and F. Markowetz, “Cancer evolution: mathematical models and computational inference,” Systematic Biology, vol. 64, no. 1, pp. e1–e25, 2014
2014
-
[103]
Person- alizing cancer therapy via machine learning,
M. Tenenbaum, A. Fern, L. Getoor, M. Littman, V . Manasinghka, S. Natarajan, D. Page, J. Shrager, Y . Singer, and P. Tadepalli, “Person- alizing cancer therapy via machine learning,” in Workshops of NIPS , 2010
2010
-
[104]
Support vector method for function approximation, regression estimation and signal processing,
V . Vapnik, S. E. Golowich, and A. J. Smola, “Support vector method for function approximation, regression estimation and signal processing,” in Advances in Neural Information Processing Systems, 1997, pp. 281– 287
1997
-
[105]
The dynamics of an optimally controlled tumor model: A case study,
L. G. De Pillis and A. Radunskaya, “The dynamics of an optimally controlled tumor model: A case study,” Mathematical and Computer Modelling, vol. 37, no. 11, pp. 1221–1244, 2003
2003
-
[106]
Machine learning in radiation oncology: Opportunities, requirements, and needs,
M. Feng, G. Valdes, N. Dixit, and T. D. Solberg, “Machine learning in radiation oncology: Opportunities, requirements, and needs,” Frontiers in Oncology, vol. 8, 2018
2018
-
[107]
Robust high performance reinforce- ment learning through weighted k-nearest neighbors,
J. de Lope, D. Maravall et al. , “Robust high performance reinforce- ment learning through weighted k-nearest neighbors,”Neurocomputing, vol. 74, no. 8, pp. 1251–1259, 2011
2011
-
[108]
Idf diabetes atlas: Global estimates of diabetes prevalence for 2017 and projections for 2045,
N. Cho, J. Shaw, S. Karuranga, Y . Huang, J. da Rocha Fernandes, A. Ohlrogge, and B. Malanda, “Idf diabetes atlas: Global estimates of diabetes prevalence for 2017 and projections for 2045,” Diabetes Research and Clinical Practice , vol. 138, pp. 271–281, 2018
2017
-
[109]
Clinical control of diabetes by the artificial pancreas,
A. M. Albisser, B. Leibel, T. Ewart, Z. Davidovac, C. Botz, W. Zingg, H. Schipper, and R. Gander, “Clinical control of diabetes by the artificial pancreas,” Diabetes, vol. 23, no. 5, pp. 397–404, 1974
1974
-
[110]
Artificial pancreas: past, present, future,
C. Cobelli, E. Renard, and B. Kovatchev, “Artificial pancreas: past, present, future,” Diabetes, vol. 60, no. 11, pp. 2672–2682, 2011
2011
-
[111]
A critical assessment of algorithms and challenges in the development of a closed-loop artificial pancreas,
B. W. Bequette, “A critical assessment of algorithms and challenges in the development of a closed-loop artificial pancreas,” Diabetes Technology & Therapeutics, vol. 7, no. 1, pp. 28–47, 2005
2005
-
[112]
The artificial pancreas: current status and future prospects in the management of diabetes,
T. Peyser, E. Dassau, M. Breton, and J. S. Skyler, “The artificial pancreas: current status and future prospects in the management of diabetes,” Annals of the New York Academy of Sciences , vol. 1311, no. 1, pp. 102–123, 2014
2014
-
[113]
The use of reinforcement learning algorithms to meet the challenges of an artificial pancreas,
M. K. Bothe, L. Dickens, K. Reichel, A. Tellmann, B. Ellger, M. West- phal, and A. A. Faisal, “The use of reinforcement learning algorithms to meet the challenges of an artificial pancreas,” Expert Review of Medical Devices, vol. 10, no. 5, pp. 661–673, 2013
2013
-
[114]
Agent-based sim- ulation for blood glucose,
S. Yasini, M. B. Naghibi Sistani, and A. Karimpour, “Agent-based sim- ulation for blood glucose,” International Journal of Applied Science, Engineering and Technology, vol. 5, pp. 89–95, 2009
2009
-
[115]
Pre- liminary results of a novel approach for glucose regulation using an actor-critic learning based controller,
E. Daskalaki, L. Scarnato, P. Diem, and S. G. Mougiakakou, “Pre- liminary results of a novel approach for glucose regulation using an actor-critic learning based controller,” 2010
2010
-
[116]
In silico preclinical trials: a proof of concept in closed-loop control of type 1 diabetes,
B. P. Kovatchev, M. Breton, C. Dalla Man, and C. Cobelli, “In silico preclinical trials: a proof of concept in closed-loop control of type 1 diabetes,” 2009
2009
-
[117]
An actor–critic based controller for glucose regulation in type 1 diabetes,
E. Daskalaki, P. Diem, and S. G. Mougiakakou, “An actor–critic based controller for glucose regulation in type 1 diabetes,”Computer Methods and Programs in Biomedicine , vol. 109, no. 2, pp. 116–125, 2013
2013
-
[118]
Personalized tuning of a reinforcement learning control al- gorithm for glucose regulation,
——, “Personalized tuning of a reinforcement learning control al- gorithm for glucose regulation,” in 2013 35th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). IEEE, 2013, pp. 3487–3490
2013
-
[119]
Model-free machine learning in biomedicine: Feasibility study in type 1 diabetes,
——, “Model-free machine learning in biomedicine: Feasibility study in type 1 diabetes,” PloS One, vol. 11, no. 7, p. e0158722, 2016
2016
-
[120]
A dual mode adaptive basal-bolus advisor based on reinforcement learning,
Q. Sun, M. Jankovic, J. Budzinski, B. Moore, P. Diem, C. Stettler, and S. G. Mougiakakou, “A dual mode adaptive basal-bolus advisor based on reinforcement learning,” IEEE journal of biomedical and health informatics, 2018
2018
-
[121]
Qualitative behavior of a family of delay-differential models of the glucose-insulin system,
P. Palumbo, S. Panunzi, and A. De Gaetano, “Qualitative behavior of a family of delay-differential models of the glucose-insulin system,” Discrete and Continuous Dynamical Systems Series B , vol. 7, no. 2, p. 399, 2007
2007
-
[122]
Glucose level control using temporal difference methods,
A. Noori, M. A. Sadrnia et al., “Glucose level control using temporal difference methods,” in 2017 Iranian Conference on Electrical Engi- neering (ICEE). IEEE, 2017, pp. 895–900
2017
-
[123]
Reinforcement-learning optimal control for type-1 diabetes,
P. D. Ngo, S. Wei, A. Holubov ´a, J. Muzik, and F. Godtliebsen, “Reinforcement-learning optimal control for type-1 diabetes,” in 2018 IEEE EMBS International Conference on Biomedical & Health Infor- matics (BHI). IEEE, 2018, pp. 333–336
2018
-
[124]
Control of blood glucose for type-1 diabetes by using rein- forcement learning with feedforward algorithm,
——, “Control of blood glucose for type-1 diabetes by using rein- forcement learning with feedforward algorithm,” Computational and Mathematical Methods in Medicine , vol. 2018, 2018
2018
-
[125]
Quantitative estimation of insulin sensitivity
R. N. Bergman, Y . Z. Ider, C. R. Bowden, and C. Cobelli, “Quantitative estimation of insulin sensitivity.” American Journal of Physiology- Endocrinology And Metabolism , vol. 236, no. 6, p. E667, 1979
1979
-
[126]
Nonlinear model predictive control of glucose concen- tration in subjects with type 1 diabetes,
R. Hovorka, V . Canonico, L. J. Chassin, U. Haueter, M. Massi- Benedetti, M. O. Federici, T. R. Pieber, H. C. Schaller, L. Schaupp, T. Veringet al., “Nonlinear model predictive control of glucose concen- tration in subjects with type 1 diabetes,” Physiological Measurement, vol...
2004
-
[127]
Controlling blood glucose variability under uncertainty using reinforcement learning and gaussian processes,
M. De Paula, L. O. ´Avila, and E. C. Mart ´ınez, “Controlling blood glucose variability under uncertainty using reinforcement learning and gaussian processes,” Applied Soft Computing , vol. 35, pp. 310–332, 2015
2015
-
[128]
On-line policy learning and adaptation for real-time personalization of an artificial pancreas,
M. De Paula, G. G. Acosta, and E. C. Mart ´ınez, “On-line policy learning and adaptation for real-time personalization of an artificial pancreas,” Expert Systems with Applications , vol. 42, no. 4, pp. 2234– 2255, 2015
2015
-
[129]
Blood glucose regulation with stochastic optimal control for insulin-dependent diabetic patients,
S. U. Acikgoz and U. M. Diwekar, “Blood glucose regulation with stochastic optimal control for insulin-dependent diabetic patients,” Chemical Engineering Science , vol. 65, no. 3, pp. 1227–1236, 2010
2010
-
[130]
Modeling medical records of diabetes using markov decision processes,
H. Asoh, M. Shiro, S. Akaho, T. Kamishima, K. Hashida, E. Aramaki, and T. Kohro, “Modeling medical records of diabetes using markov decision processes,” in Proceedings of ICML2013 Workshop on Role of Machine Learning in Transforming Healthcare , 2013
2013
-
[131]
An application of inverse reinforcement learning to medical records of diabetes treatment,
H. Asoh, M. S. S. Akaho, T. Kamishima, K. Hasida, E. Aramaki, and T. Kohro, “An application of inverse reinforcement learning to medical records of diabetes treatment,” in ECMLPKDD2013 Workshop on Reinforcement Learning with Generalized Feedback , 2013
2013
-
[132]
Estimating dynamic treatment regimes in mobile health using v-learning,
D. J. Luckett, E. B. Laber, A. R. Kahkoska, D. M. Maahs, E. Mayer- Davis, and M. R. Kosorok, “Estimating dynamic treatment regimes in mobile health using v-learning,” Journal of the American Statistical Association, no. just-accepted, pp. 1–39, 2018
2018
-
[133]
Reinforcement learning approach to individualization of chronic pharmacotherapy,
A. E. Gaweda, M. K. Muezzinoglu, G. R. Aronoff, A. A. Jacobs, J. M. Zurada, and M. E. Brier, “Reinforcement learning approach to individualization of chronic pharmacotherapy,” in IJCNN’05, vol. 5. IEEE, 2005, pp. 3290–3295
2005
-
[134]
Model predictive control with reinforcement learning for drug delivery in renal anemia management,
A. E. Gaweda, M. K. Muezzinoglu, A. A. Jacobs, G. R. Aronoff, and M. E. Brier, “Model predictive control with reinforcement learning for drug delivery in renal anemia management,” in IEEE EMBS’06. IEEE, 2006, pp. 5177–5180
2006
-
[135]
Individualization of pharmacological anemia management using reinforcement learning,
A. E. Gaweda, M. K. Muezzinoglu, G. R. Aronoff, A. A. Jacobs, J. M. Zurada, and M. E. Brier, “Individualization of pharmacological anemia management using reinforcement learning,” Neural Networks, vol. 18, no. 5-6, pp. 826–834, 2005
2005
-
[136]
A reinforcement learn- ing approach for individualizing erythropoietin dosages in hemodialysis patients,
J. D. Mart ´ın-Guerrero, F. Gomez, E. Soria-Olivas, J. Schmidhuber, M. Climente-Mart´ı, and N. V . Jim´enez-Torres, “A reinforcement learn- ing approach for individualizing erythropoietin dosages in hemodialysis patients,” Expert Systems with Applications , vol. 36, no. 6, pp....
2009
-
[137]
Validation of a reinforcement learning policy for dosage optimization of erythropoietin,
J. D. Mart ´ın-Guerrero, E. Soria-Olivas, M. Mart ´ınez-Sober, M. Climente-Mart ´ı, T. De Diego-Santos, and N. V . Jim ´enez- Torres, “Validation of a reinforcement learning policy for dosage optimization of erythropoietin,” in Australasian Joint Conference on Artificial Intell...
2007
-
[138]
Optimizing drug therapy with reinforcement learning: The case of anemia management,
J. M. Malof and A. E. Gaweda, “Optimizing drug therapy with reinforcement learning: The case of anemia management,” in Neural Networks (IJCNN), The 2011 International Joint Conference on. IEEE, 2011, pp. 2088–2092
2011
-
[139]
Adaptive treatment of anemia on hemodialysis patients: A reinforcement learning approach,
P. Escandell-Montero, J. M. Mart ´ınez-Mart´ınez, J. D. Mart´ın-Guerrero, E. Soria-Olivas, J. Vila-Franc´es, and R. Magdalena-Benedito, “Adaptive treatment of anemia on hemodialysis patients: A reinforcement learning approach,” in CIDM2011. IEEE, 2011, pp. 44–49
2011
-
[140]
Optimization of anemia treatment in hemodialysis patients via reinforcement learning,
P. Escandell-Montero, M. Chermisi, J. M. Martinez-Martinez, J. Gomez-Sanchis, C. Barbieri, E. Soria-Olivas, F. Mari, J. Vila- Franc´es, A. Stopper, E. Gatti et al., “Optimization of anemia treatment in hemodialysis patients via reinforcement learning,” Artificial Intelli- gence...
2014
-
[141]
Dynamic multidrug therapies for hiv: Optimal and sti control approaches,
B. M. Adams, H. T. Banks, H.-D. Kwon, and H. T. Tran, “Dynamic multidrug therapies for hiv: Optimal and sti control approaches,” Mathematical Biosciences and Engineering, vol. 1, no. 2, pp. 223–241, 2004
2004
-
[142]
Clinical data based optimal sti strategies for hiv: a reinforcement learning approach,
D. Ernst, G.-B. Stan, J. Goncalves, and L. Wehenkel, “Clinical data based optimal sti strategies for hiv: a reinforcement learning approach,” in 45th IEEE Conference on Decision and Control . IEEE, 2006, pp. 667–672
2006
-
[143]
A reinforcement learning design for hiv clinical trials,
S. Parbhoo, “A reinforcement learning design for hiv clinical trials,” Ph.D. dissertation, 2014
2014
-
[144]
Combining kernel and model based learning for hiv therapy selection,
S. Parbhoo, J. Bogojeska, M. Zazzi, V . Roth, and F. Doshi-Velez, “Combining kernel and model based learning for hiv therapy selection,” AMIA Summits on Translational Science Proceedings , vol. 2017, p. 239, 2017
2017
-
[145]
Quanti- fying uncertainty in batch personalized sequential decision making
V . N. Marivate, J. Chemali, E. Brunskill, and M. L. Littman, “Quanti- fying uncertainty in batch personalized sequential decision making.” in AAAI Workshop: Modern Artificial Intelligence for Health Analytics , 2014
2014
-
[146]
Transfer learning across patient variations with hidden parameter markov decision processes,
T. Killian, G. Konidaris, and F. Doshi-Velez, “Transfer learning across patient variations with hidden parameter markov decision processes,” arXiv preprint arXiv:1612.00475 , 2016
2016 arXiv
-
[147]
Robust and efficient transfer learning with hidden parameter markov decision processes,
T. W. Killian, S. Daulton, G. Konidaris, and F. Doshi-Velez, “Robust and efficient transfer learning with hidden parameter markov decision processes,” in Advances in Neural Information Processing Systems , 2017, pp. 6250–6261
2017
-
[148]
Direct policy transfer via hidden parameter markov decision processes,
J. Yao, T. Killian, G. Konidaris, and F. Doshi-Velez, “Direct policy transfer via hidden parameter markov decision processes,” 2018
2018
-
[149]
Incorporating causal factors into reinforcement learning for dynamic treatment regimes in hiv,
C. Yu, Y . Dong, J. Liu, and G. Ren, “Incorporating causal factors into reinforcement learning for dynamic treatment regimes in hiv,” BMC medical informatics and decision making , vol. 19, no. 2, p. 60, 2019
2019
-
[150]
Pac optimal exploration in continuous space markov decision processes
J. Pazis and R. Parr, “Pac optimal exploration in continuous space markov decision processes.” in AAAI, 2013
2013
-
[151]
Bounded optimal exploration in mdp
K. Kawaguchi, “Bounded optimal exploration in mdp.” in AAAI, 2016, pp. 1758–1764
2016
-
[152]
Methodological challenges in constructing effective treatment sequences for chronic psychiatric disorders,
S. A. Murphy, D. W. Oslin, A. J. Rush, and J. Zhu, “Methodological challenges in constructing effective treatment sequences for chronic psychiatric disorders,” Neuropsychopharmacology, vol. 32, no. 2, p. 257, 2007
2007
-
[153]
Eeg seizure detection and prediction algorithms: a survey,
T. N. Alotaiby, S. A. Alshebeili, T. Alshawi, I. Ahmad, and F. E. A. El-Samie, “Eeg seizure detection and prediction algorithms: a survey,” EURASIP Journal on Advances in Signal Processing , vol. 2014, no. 1, p. 183, 2014
2014
-
[154]
Progress in neuroengineering for brain repair: New challenges and open issues,
G. Panuccio, M. Semprini, L. Natale, S. Buccelli, I. Colombi, and M. Chiappalone, “Progress in neuroengineering for brain repair: New challenges and open issues,” Brain and Neuroscience Advances, vol. 2, p. 2398212818776475, 2018
2018
-
[155]
Adaptive treatment of epilepsy via batch-mode reinforcement learning
A. Guez, R. D. Vincent, M. Avoli, and J. Pineau, “Adaptive treatment of epilepsy via batch-mode reinforcement learning.” in AAAI, 2008, pp. 1671–1678
2008
-
[156]
Treating epilepsy via adaptive neurostimulation: a reinforcement learning ap- proach,
J. Pineau, A. Guez, R. Vincent, G. Panuccio, and M. Avoli, “Treating epilepsy via adaptive neurostimulation: a reinforcement learning ap- proach,” International Journal of Neural Systems , vol. 19, no. 04, pp. 227–240, 2009
2009
-
[157]
Adaptive control of epileptic seizures using reinforcement learning,
A. Guez, “Adaptive control of epileptic seizures using reinforcement learning,” Ph.D. dissertation, McGill University Library, 2010
2010
-
[158]
Adaptive control of epileptiform excitability in an in vitro model of limbic seizures,
G. Panuccio, A. Guez, R. Vincent, M. Avoli, and J. Pineau, “Adaptive control of epileptiform excitability in an in vitro model of limbic seizures,” Experimental Neurology, vol. 241, pp. 179–183, 2013
2013
-
[159]
Manifold embeddings for model-based rein- forcement learning under partial observability,
K. Bush and J. Pineau, “Manifold embeddings for model-based rein- forcement learning under partial observability,” in Advances in Neural Information Processing Systems , 2009, pp. 189–197
2009
-
[160]
Seizure control in a computational model using a reinforcement learning stimulation paradigm,
V . Nagaraj, A. Lamperski, and T. I. Netoff, “Seizure control in a computational model using a reinforcement learning stimulation paradigm,” International Journal of Neural Systems , vol. 27, no. 07, p. 1750012, 2017
2017
-
[161]
Sequenced treatment alternatives to relieve depression (star* d): rationale and design,
A. J. Rush, M. Fava, S. R. Wisniewski, P. W. Lavori, M. H. Trivedi, H. A. Sackeim, M. E. Thase, A. A. Nierenberg, F. M. Quitkin, T. M. Kashner et al., “Sequenced treatment alternatives to relieve depression (star* d): rationale and design,”Controlled clinical trials, vol. 25, ...
2004
-
[162]
Constructing evidence-based treatment strategies using methods from computer science,
J. Pineau, M. G. Bellemare, A. J. Rush, A. Ghizaru, and S. A. Murphy, “Constructing evidence-based treatment strategies using methods from computer science,” Drug & Alcohol Dependence, vol. 88, pp. S52–S60, 2007
2007
-
[163]
Kernel-based reinforcement learning,
D. Ormoneit and ´S. Sen, “Kernel-based reinforcement learning,” Ma- chine learning, vol. 49, no. 2-3, pp. 161–178, 2002
2002
-
[164]
Inference for optimal dynamic treatment regimes using an adaptive m-out-of-n bootstrap scheme,
B. Chakraborty, E. B. Laber, and Y . Zhao, “Inference for optimal dynamic treatment regimes using an adaptive m-out-of-n bootstrap scheme,” Biometrics, vol. 69, no. 3, pp. 714–723, 2013
2013
-
[165]
Interactive model building for q-learning,
E. B. Laber, K. A. Linn, and L. A. Stefanski, “Interactive model building for q-learning,” Biometrika, vol. 101, no. 4, pp. 831–847, 2014
2014
-
[166]
Interactive q-learning for probabilities and quantiles,
K. A. Linn, E. B. Laber, and L. A. Stefanski, “Interactive q-learning for probabilities and quantiles,” arXiv preprint arXiv:1407.3414, 2014
2014 arXiv
-
[167]
Interactive q-learning for quantiles,
——, “Interactive q-learning for quantiles,” Journal of the American Statistical Association, vol. 112, no. 518, pp. 638–649, 2017
2017
-
[168]
Q-and a- learning methods for estimating optimal dynamic treatment regimes,
P. J. Schulte, A. A. Tsiatis, E. B. Laber, and M. Davidian, “Q-and a- learning methods for estimating optimal dynamic treatment regimes,” Statistical science: a review journal of the Institute of Mathematical Statistics, vol. 29, no. 4, p. 640, 2014
2014
-
[169]
Optimal dynamic treatment regimes,
S. A. Murphy, “Optimal dynamic treatment regimes,” Journal of the Royal Statistical Society: Series B (Statistical Methodology) , vol. 65, no. 2, pp. 331–355, 2003
2003
-
[170]
Penalized q-learning for dynamic treatment regimens,
R. Song, W. Wang, D. Zeng, and M. R. Kosorok, “Penalized q-learning for dynamic treatment regimens,” Statistica Sinica , vol. 25, no. 3, p. 901, 2015
2015
-
[171]
Robust hybrid learning for estimating personalized dynamic treatment regimens,
Y . Liu, Y . Wang, M. R. Kosorok, Y . Zhao, and D. Zeng, “Robust hybrid learning for estimating personalized dynamic treatment regimens,” arXiv preprint arXiv:1611.02314 , 2016
2016 arXiv
-
[172]
Budgeted learning for developing personalized treatment,
K. Deng, R. Greiner, and S. Murphy, “Budgeted learning for developing personalized treatment,” in ICMLA2014. IEEE, 2014, pp. 7–14
2014
-
[173]
Neurocognitive effects of antipsychotic medications in patients with chronic schizophrenia in the catie trial,
R. S. Keefe, R. M. Bilder, S. M. Davis, P. D. Harvey, B. W. Palmer, J. M. Gold, H. Y . Meltzer, M. F. Green, G. Capuano, T. S. Stroup et al., “Neurocognitive effects of antipsychotic medications in patients with chronic schizophrenia in the catie trial,” Archives of General Ps...
2007
-
[174]
Informing sequential clinical decision-making through reinforcement learning: an empirical study,
S. M. Shortreed, E. Laber, D. J. Lizotte, T. S. Stroup, J. Pineau, and S. A. Murphy, “Informing sequential clinical decision-making through reinforcement learning: an empirical study,”Machine Learning, vol. 84, no. 1-2, pp. 109–136, 2011
2011
-
[175]
Q-learning residual analysis: application to the effectiveness of sequences of antipsychotic medications for patients with schizophrenia,
A. Ertefaie, S. Shortreed, and B. Chakraborty, “Q-learning residual analysis: application to the effectiveness of sequences of antipsychotic medications for patients with schizophrenia,” Statistics in Medicine , vol. 35, no. 13, pp. 2221–2234, 2016
2016
-
[176]
Linear fitted-q iter- ation with multiple reward functions,
D. J. Lizotte, M. Bowling, and S. A. Murphy, “Linear fitted-q iter- ation with multiple reward functions,” Journal of Machine Learning Research, vol. 13, no. Nov, pp. 3253–3295, 2012
2012
-
[177]
Multi-objective markov decision processes for data-driven decision support,
D. J. Lizotte and E. B. Laber, “Multi-objective markov decision processes for data-driven decision support,” The Journal of Machine Learning Research, vol. 17, no. 1, pp. 7378–7405, 2016
2016
-
[178]
Set-valued dynamic treatment regimes for competing outcomes,
E. B. Laber, D. J. Lizotte, and B. Ferguson, “Set-valued dynamic treatment regimes for competing outcomes,” Biometrics, vol. 70, no. 1, pp. 53–61, 2014
2014
-
[179]
Incor- porating patient preferences into estimation of optimal individualized treatment rules,
E. L. Butler, E. B. Laber, S. M. Davis, and M. R. Kosorok, “Incor- porating patient preferences into estimation of optimal individualized treatment rules,” Biometrics, 2017
2017
-
[180]
Managing addiction as a chronic con- dition,
M. Dennis and C. K. Scott, “Managing addiction as a chronic con- dition,” Addiction Science & Clinical Practice , vol. 4, no. 1, p. 45, 2007
2007
-
[181]
A batch, off-policy, actor-critic algorithm for optimizing the average reward,
S. A. Murphy, Y . Deng, E. B. Laber, H. R. Maei, R. S. Sutton, and K. Witkiewitz, “A batch, off-policy, actor-critic algorithm for optimizing the average reward,” arXiv preprint arXiv:1607.05047 , 2016
2016 arXiv
-
[182]
Inference for non-regular parameters in optimal dynamic treatment regimes,
B. Chakraborty, S. Murphy, and V . Strecher, “Inference for non-regular parameters in optimal dynamic treatment regimes,” Statistical Methods in Medical Research , vol. 19, no. 3, pp. 317–343, 2010
2010
-
[183]
Bias correction and confidence intervals for fitted q-iteration,
B. Chakraborty, V . Strecher, and S. Murphy, “Bias correction and confidence intervals for fitted q-iteration,” in Workshop on Model Uncertainty and Risk in Reinforcement Learning, NIPS, Whistler, Canada. Citeseer, 2008
2008
-
[184]
Tree-based reinforcement learning for estimating optimal dynamic treatment regimes,
Y . Tao, L. Wang, D. Almirall et al., “Tree-based reinforcement learning for estimating optimal dynamic treatment regimes,” The Annals of Applied Statistics, vol. 12, no. 3, pp. 1914–1938, 2018
1914
-
[185]
Critical care-where have we been and where are we going?
J.-L. Vincent, “Critical care-where have we been and where are we going?” Critical Care, vol. 17, no. 1, p. S2, 2013
2013
-
[186]
Critical care workforce,
K. Krell, “Critical care workforce,” Critical Care Medicine , vol. 36, no. 4, pp. 1350–1353, 2008
2008
-
[187]
State of the art review: the data revolution in critical care,
M. Ghassemi, L. A. Celi, and D. J. Stone, “State of the art review: the data revolution in critical care,” Critical Care, vol. 19, no. 1, p. 118, 2015
2015
-
[188]
Surviving sepsis campaign: international guidelines for man- agement of sepsis and septic shock: 2016,
A. Rhodes, L. E. Evans, W. Alhazzani, M. M. Levy, M. Antonelli, R. Ferrer, A. Kumar, J. E. Sevransky, C. L. Sprung, M. E. Nunnally et al. , “Surviving sepsis campaign: international guidelines for man- agement of sepsis and septic shock: 2016,” Intensive Care Medicine , vol. 4...
2016
-
[189]
Acute respiratory distress syndrome,
A. D. T. Force, V . Ranieri, G. Rubenfeld et al. , “Acute respiratory distress syndrome,” Jama, vol. 307, no. 23, pp. 2526–2533, 2012
2012
-
[190]
Use of machine-learning approaches to predict clinical deterioration in critically ill patients: A systematic review,
T. Kamio, T. Van, and K. Masamune, “Use of machine-learning approaches to predict clinical deterioration in critically ill patients: A systematic review,” International Journal of Medical Research and Health Sciences, vol. 6, no. 6, pp. 1–7, 2017
2017
-
[191]
Machine learning in critical care: state-of-the-art and a sepsis case study,
A. Vellido, V . Ribas, C. Morales, A. R. Sanmart ´ın, and J. C. R. Rodr´ıguez, “Machine learning in critical care: state-of-the-art and a sepsis case study,” Biomedical engineering online , vol. 17, no. 1, p. 135, 2018
2018
-
[192]
A markov decision process to suggest optimal treatment of severe infections in intensive care,
M. Komorowski, A. Gordon, L. Celi, and A. Faisal, “A markov decision process to suggest optimal treatment of severe infections in intensive care,” in Neural Information Processing Systems Workshop on Machine Learning for Health , 2016
2016
-
[193]
The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care,
M. Komorowski, L. A. Celi, O. Badawi, A. C. Gordon, and A. A. Faisal, “The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care,”Nature Medicine, vol. 24, no. 11, p. 1716, 2018
2018
-
[194]
Deep reinforcement learning for sepsis treatment,
A. Raghu, M. Komorowski, I. Ahmed, L. Celi, P. Szolovits, and M. Ghassemi, “Deep reinforcement learning for sepsis treatment,” arXiv preprint arXiv:1711.09602 , 2017
2017 arXiv
-
[195]
Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach,
A. Raghu, M. Komorowski, L. A. Celi, P. Szolovits, and M. Ghassemi, “Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach,” in Machine Learning for Healthcare Conference, 2017, pp. 147–163
2017
-
[196]
Model-based reinforcement learning for sepsis treatment,
A. Raghu, M. Komorowski, and S. Singh, “Model-based reinforcement learning for sepsis treatment,” arXiv preprint arXiv:1811.09602, 2018
2018 arXiv
-
[197]
Treatment recommendation in crit- ical care: A scalable and interpretable approach in partially observable health states,
C. P. Utomo, X. Li, and W. Chen, “Treatment recommendation in crit- ical care: A scalable and interpretable approach in partially observable health states,” 2018
2018
-
[198]
Improving sepsis treatment strategies by combining deep and kernel-based reinforcement learning,
X. Peng, Y . Ding, D. Wihl, O. Gottesman, M. Komorowski, L.-w. H. Lehman, A. Ross, A. Faisal, and F. Doshi-Velez, “Improving sepsis treatment strategies by combining deep and kernel-based reinforcement learning,” arXiv preprint arXiv:1901.04670 , 2019
1901 arXiv
-
[199]
Learning to treat sepsis with multi-output gaussian process deep recurrent q-networks,
J. Futoma, A. Lin, M. Sendak, A. Bedoya, M. Clement, C. O’Brien, and K. Heller, “Learning to treat sepsis with multi-output gaussian process deep recurrent q-networks,” 2018
2018
-
[200]
Deep inverse reinforcement learning for sepsis treatment,
C. Yu, G. Ren, and J. Liu, “Deep inverse reinforcement learning for sepsis treatment,” in 2019 IEEE ICHI , 2019, pp. 1–3
2019
-
[201]
The actor search tree critic (astc) for off-policy pomdp learning in medical decision making,
L. Li, M. Komorowski, and A. A. Faisal, “The actor search tree critic (astc) for off-policy pomdp learning in medical decision making,”arXiv preprint arXiv:1805.11548, 2018
2018 arXiv
-
[202]
Representation and reinforcement learning for personalized glycemic control in septic patients,
W.-H. Weng, M. Gao, Z. He, S. Yan, and P. Szolovits, “Representation and reinforcement learning for personalized glycemic control in septic patients,” arXiv preprint arXiv:1712.00654 , 2017
2017 arXiv
-
[203]
Precision medicine as a control problem: Using simulation and deep reinforcement learning to discover adap- tive, personalized multi-cytokine therapy for sepsis,
B. K. Petersen, J. Yang, W. S. Grathwohl, C. Cockrell, C. Santiago, G. An, and D. M. Faissol, “Precision medicine as a control problem: Using simulation and deep reinforcement learning to discover adap- tive, personalized multi-cytokine therapy for sepsis,” arXiv preprint arXi...
2018 arXiv
-
[204]
Intelligent control of closed-loop sedation in simulated icu patients
B. L. Moore, E. D. Sinzinger, T. M. Quasny, and L. D. Pyeatt, “Intelligent control of closed-loop sedation in simulated icu patients.” in FLAIRS Conference, 2004, pp. 109–114
2004
-
[205]
Sedation of simulated icu patients using reinforcement learning based control,
E. D. Sinzinger and B. Moore, “Sedation of simulated icu patients using reinforcement learning based control,” International Journal on Artificial Intelligence Tools, vol. 14, no. 01n02, pp. 137–156, 2005
2005
-
[206]
Reinforcement learning: a novel method for optimal control of propofol-induced hypnosis,
B. L. Moore, A. G. Doufas, and L. D. Pyeatt, “Reinforcement learning: a novel method for optimal control of propofol-induced hypnosis,” Anesthesia & Analgesia , vol. 112, no. 2, pp. 360–367, 2011
2011
-
[207]
Reinforcement learning versus proportional–integral–derivative control of hypnosis in a simulated intraoperative patient,
B. L. Moore, T. M. Quasny, and A. G. Doufas, “Reinforcement learning versus proportional–integral–derivative control of hypnosis in a simulated intraoperative patient,” Anesthesia & Analgesia , vol. 112, no. 2, pp. 350–359, 2011
2011
-
[208]
Reinforcement learning for closed-loop propofol anesthesia: a study in human volunteers,
B. L. Moore, L. D. Pyeatt, V . Kulkarni, P. Panousis, K. Padrez, and A. G. Doufas, “Reinforcement learning for closed-loop propofol anesthesia: a study in human volunteers,” The Journal of Machine Learning Research, vol. 15, no. 1, pp. 655–696, 2014
2014
-
[209]
Reinforcement learning for closed-loop propofol anesthesia: A human volunteer study
B. L. Moore, P. Panousis, V . Kulkarni, L. D. Pyeatt, and A. G. Doufas, “Reinforcement learning for closed-loop propofol anesthesia: A human volunteer study.” in IAAI, 2010
2010
-
[210]
Multivariable anesthesia control using reinforcement learning,
N. Sadati, A. Aflaki, and M. Jahed, “Multivariable anesthesia control using reinforcement learning,” in IEEE SMC’06, vol. 6. IEEE, 2006, pp. 4563–4568
2006
-
[211]
An adaptive neural network filter for improved patient state estimation in closed- loop anesthesia control,
E. C. Borera, B. L. Moore, A. G. Doufas, and L. D. Pyeatt, “An adaptive neural network filter for improved patient state estimation in closed- loop anesthesia control,” in IEEE ICTAI’11. IEEE, 2011, pp. 41–46
2011
-
[212]
Towards efficient, personalized anesthesia using continuous reinforcement learning for propofol infusion control,
C. Lowery and A. A. Faisal, “Towards efficient, personalized anesthesia using continuous reinforcement learning for propofol infusion control,” in IEEE/EMBS NER’13. IEEE, 2013, pp. 1414–1417
2013
-
[213]
Closed-loop control of anesthesia and mean arterial pressure using reinforcement learning,
R. Padmanabhan, N. Meskin, and W. M. Haddad, “Closed-loop control of anesthesia and mean arterial pressure using reinforcement learning,” Biomedical Signal Processing and Control , vol. 22, pp. 54–64, 2015
2015
-
[214]
Learning from an expert
P. Humbert, J. Audiffren, C. Dubost, and L. Oudre, “Learning from an expert.”
-
[215]
Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach,
S. Nemati, M. M. Ghassemi, and G. D. Clifford, “Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach,” in IEEE 38th Annual International Conference of the Engineering in Medicine and Biology Society . IEEE, 2016, pp. 2978–2981
2016
-
[216]
A deep deter- ministic policy gradient approach to medication dosing and surveillance in the icu,
R. Lin, M. D. Stanley, M. M. Ghassemi, and S. Nemati, “A deep deter- ministic policy gradient approach to medication dosing and surveillance in the icu,” in IEEE EMBC’18. IEEE, 2018, pp. 4927–4931
2018
-
[217]
Supervised reinforcement learning with recurrent neural network for dynamic treatment recom- mendation,
L. Wang, W. Zhang, X. He, and H. Zha, “Supervised reinforcement learning with recurrent neural network for dynamic treatment recom- mendation,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . ACM, 2018, pp. 2447–2456
2018
-
[218]
A reinforcement learning approach to weaning of mechanical ventilation in intensive care units,
N. Prasad, L.-F. Cheng, C. Chivers, M. Draugelis, and B. E. Engel- hardt, “A reinforcement learning approach to weaning of mechanical ventilation in intensive care units,” arXiv preprint arXiv:1704.06300 , 2017
2017 arXiv
-
[219]
Inverse reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units,
C. Yu, J. Liu, and H. Zhao, “Inverse reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units,” BMC medical informatics and decision making , vol. 19, no. 2, p. 57, 2019
2019
-
[220]
Supervised-actor-critic reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units,
C. Yu, G. Ren, and Y . Dong, “Supervised-actor-critic reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units,” BMC medical informatics and decision making , 2020
2020
-
[221]
Towards high confidence off- policy reinforcement learning for clinical applications
A. Jagannatha, P. Thomas, and H. Yu, “Towards high confidence off- policy reinforcement learning for clinical applications.”
-
[222]
An optimal policy for patient laboratory tests in intensive care units,
L.-F. Cheng, N. Prasad, and B. E. Engelhardt, “An optimal policy for patient laboratory tests in intensive care units,” arXiv preprint arXiv:1808.04679, 2018
2018 arXiv
-
[223]
Dynamic measurement scheduling for adverse event forecasting using deep rl,
C.-H. Chang, M. Mai, and A. Goldenberg, “Dynamic measurement scheduling for adverse event forecasting using deep rl,” arXiv preprint arXiv:1812.00268, 2018
2018 arXiv
-
[224]
Tools for the precision medicine era: How to develop highly personalized treatment recommendations from cohort and registry data using q-learning,
E. F. Krakow, M. Hemmer, T. Wang, B. Logan, M. Arora, S. Spellman, D. Couriel, A. Alousi, J. Pidala, M. Last et al. , “Tools for the precision medicine era: How to develop highly personalized treatment recommendations from cohort and registry data using q-learning,” American j...
2017
-
[225]
Deep reinforcement learning for dynamic treatment regimes on medical registry data,
Y . Liu, B. Logan, N. Liu, Z. Xu, J. Tang, and Y . Wang, “Deep reinforcement learning for dynamic treatment regimes on medical registry data,” in IEEE ICHI’17. IEEE, 2017, pp. 380–385
2017
-
[226]
Sepsis: pathophysiology and clinical management,
J. E. Gotts and M. A. Matthay, “Sepsis: pathophysiology and clinical management,” Bmj, vol. 353, p. i1585, 2016
2016
-
[227]
Mimic-iii, a freely accessible critical care database,
A. E. Johnson, T. J. Pollard, L. Shen, H. L. Li-wei, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark, “Mimic-iii, a freely accessible critical care database,” Scientific Data, vol. 3, p. 160035, 2016
2016
-
[228]
Individualized sepsis treatment using reinforcement learn- ing,
S. Saria, “Individualized sepsis treatment using reinforcement learn- ing,” Nature medicine, vol. 24, no. 11, p. 1641, 2018
2018
-
[229]
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning.” in AAAI, vol. 2. Phoenix, AZ, 2016, p. 5
2016
-
[230]
Dueling network architectures for deep reinforcement learning,
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in International Conference on Machine Learning , 2016, pp. 1995–2003
2016
-
[231]
Prioritized experience replay,
T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” arXiv preprint arXiv:1511.05952 , 2015
2015 arXiv
-
[232]
Prox- imal policy optimization algorithms,
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Prox- imal policy optimization algorithms,”arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[233]
Continuous control with deep reinforce- ment learning,
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y . Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforce- ment learning,” arXiv preprint arXiv:1509.02971 , 2015
2015 arXiv
-
[234]
Policy invariance under reward transformations: Theory and application to reward shaping,
A. Y . Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in ICML, vol. 99, 1999, pp. 278–287
1999
-
[235]
Clinical decision support and closed-loop control for intensive care unit sedation,
W. M. Haddad, J. M. Bailey, B. Gholami, and A. R. Tannenbaum, “Clinical decision support and closed-loop control for intensive care unit sedation,” Asian Journal of Control, vol. 20, no. 5, pp. 1343–1350, 2012
2012
-
[236]
A data-driven approach to optimized medication dosing: a focus on heparin,
M. M. Ghassemi, S. E. Richter, I. M. Eche, T. W. Chen, J. Danziger, and L. A. Celi, “A data-driven approach to optimized medication dosing: a focus on heparin,” Intensive Care Medicine , vol. 40, no. 9, pp. 1332– 1339, 2014
2014
-
[237]
The intensive care medicine research agenda for airways, invasive and noninvasive mechanical ventilation,
S. Jaber, G. Bellani, L. Blanch, A. Demoule, A. Esteban, L. Gattinoni, C. Gu ´erin, N. Hill, J. G. Laffey, S. M. Maggiore et al., “The intensive care medicine research agenda for airways, invasive and noninvasive mechanical ventilation,” Intensive Care Medicine , vol. 43, no. ...
2017
-
[238]
Focus on ventilation and airway management in the icu,
A. De Jong, G. Citerio, and S. Jaber, “Focus on ventilation and airway management in the icu,” Intensive Care Medicine , vol. 43, no. 12, pp. 1912–1915, 2017
1912
-
[239]
National Academies of Sciences, Medicine et al
E. National Academies of Sciences, Medicine et al. , Improving diag- nosis in health care . National Academies Press, 2016
2016
-
[240]
A review on use of machine learning techniques in diagnostic health-care,
S. K. Rai and K. Sowmya, “A review on use of machine learning techniques in diagnostic health-care,” Artificial Intelligent Systems and Machine Learning, vol. 10, no. 4, pp. 102–107, 2018
2018
-
[241]
Survey of machine learning algorithms for disease diagnostic,
M. Fatima and M. Pasha, “Survey of machine learning algorithms for disease diagnostic,” Journal of Intelligent Learning Systems and Applications, vol. 9, no. 01, p. 1, 2017
2017
-
[242]
Disease diagnosis in smart healthcare: Innovation, technologies and applications,
K. T. Chui, W. Alhalabi, S. S. H. Pang, P. O. d. Pablos, R. W. Liu, and M. Zhao, “Disease diagnosis in smart healthcare: Innovation, technologies and applications,” Sustainability, vol. 9, no. 12, p. 2309, 2017
2017
-
[243]
Learning to diagnose with lstm recurrent neural networks,
Z. C. Lipton, D. C. Kale, C. Elkan, and R. Wetzel, “Learning to diagnose with lstm recurrent neural networks,” arXiv preprint arXiv:1511.03677, 2015
2015 arXiv
-
[244]
Retain: An interpretable predictive model for healthcare using reverse time attention mechanism,
E. Choi, M. T. Bahadori, J. Sun, J. Kulas, A. Schuetz, and W. Stewart, “Retain: An interpretable predictive model for healthcare using reverse time attention mechanism,” in Advances in Neural Information Pro- cessing Systems, 2016, pp. 3504–3512
2016
-
[245]
Medical question answering for clinical decision support,
T. R. Goodwin and S. M. Harabagiu, “Medical question answering for clinical decision support,” in Proceedings of the 25th ACM Inter- national on Conference on Information and Knowledge Management . ACM, 2016, pp. 297–306
2016
-
[246]
Diagnostic inferencing via improving clinical concept extraction with deep reinforcement learning: A preliminary study,
Y . Ling, S. A. Hasan, V . Datla, A. Qadir, K. Lee, J. Liu, and O. Farri, “Diagnostic inferencing via improving clinical concept extraction with deep reinforcement learning: A preliminary study,” in Machine Learn- ing for Healthcare Conference , 2017, pp. 271–285
2017
-
[247]
Reinforcement learning in computer vision,
A. Bernstein and E. Burnaev, “Reinforcement learning in computer vision,” in CMV’17, vol. 10696. International Society for Optics and Photonics, 2018, p. 106961S
2018
-
[248]
A reinforcement learning framework for parameter control in computer vision applications,
G. W. Taylor, “A reinforcement learning framework for parameter control in computer vision applications,” in Computer and Robot Vision, 2004. Proceedings. First Canadian Conference on . IEEE, 2004, pp. 496–503
2004
-
[249]
A reinforcement learning framework for medical image segmentation,
F. Sahba, H. R. Tizhoosh, and M. M. Salama, “A reinforcement learning framework for medical image segmentation,” in IJCNN, vol. 6, 2006, pp. 511–517
2006
-
[250]
Application of opposition-based reinforcement learning in image segmentation,
——, “Application of opposition-based reinforcement learning in image segmentation,” in 2007 IEEE Symposium on Computational Intelligence in Image and Signal Processing . IEEE, 2007, pp. 246– 251
2007
-
[251]
Application of reinforcement learning for segmentation of transrectal ultrasound images,
——, “Application of reinforcement learning for segmentation of transrectal ultrasound images,” BMC Medical Imaging , vol. 8, no. 1, p. 8, 2008
2008
-
[252]
Object segmentation in image sequences using reinforce- ment learning,
F. Sahba, “Object segmentation in image sequences using reinforce- ment learning,” in CSCI’16. IEEE, 2016, pp. 1416–1417
2016
-
[253]
Deep reinforcement learning for surgical gesture segmentation and classification,
D. Liu and T. Jiang, “Deep reinforcement learning for surgical gesture segmentation and classification,” in International Conference on Med- ical Image Computing and Computer-Assisted Intervention . Springer, 2018, pp. 247–255
2018
-
[254]
An artificial agent for anatomical landmark detection in medical images,
F. C. Ghesu, B. Georgescu, T. Mansi, D. Neumann, J. Hornegger, and D. Comaniciu, “An artificial agent for anatomical landmark detection in medical images,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2016, pp. 229–237
2016
-
[255]
Multi-scale deep reinforcement learning for real- time 3d-landmark detection in ct scans,
F. C. Ghesu, B. Georgescu, Y . Zheng, S. Grbic, A. Maier, J. Hornegger, and D. Comaniciu, “Multi-scale deep reinforcement learning for real- time 3d-landmark detection in ct scans,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2017
2017
-
[256]
Towards intelligent robust detection of anatomical structures in incomplete volumetric data,
F. C. Ghesu, B. Georgescu, S. Grbic, A. Maier, J. Hornegger, and D. Comaniciu, “Towards intelligent robust detection of anatomical structures in incomplete volumetric data,” Medical Image Analysis , vol. 48, pp. 203–213, 2018
2018
-
[257]
Nonlinear adaptively learned optimization for object localization in 3d medical images,
M. Etcheverry, B. Georgescu, B. Odry, T. J. Re, S. Kaushik, B. Geiger, N. Mariappan, S. Grbic, and D. Comaniciu, “Nonlinear adaptively learned optimization for object localization in 3d medical images,” in Deep Learning in Medical Image Analysis and Multimodal Learning for Cli...
2018
-
[258]
Evaluating reinforcement learning agents for anatomical landmark detection,
A. Alansary, O. Oktay, Y . Li, L. Le Folgoc, B. Hou, G. Vaillant, B. Glocker, B. Kainz, and D. Rueckert, “Evaluating reinforcement learning agents for anatomical landmark detection,” 2018
2018
-
[259]
Automatic view planning with multi-scale deep reinforcement learning agents,
A. Alansary, L. L. Folgoc, G. Vaillant, O. Oktay, Y . Li, W. Bai, J. Passerat-Palmbach, R. Guerrero, K. Kamnitsas, B. Hou et al. , “Automatic view planning with multi-scale deep reinforcement learning agents,” arXiv preprint arXiv:1806.03228 , 2018
2018 arXiv
-
[260]
Partial policy-based reinforcement learning for anatomical landmark localization in 3d medical images,
W. A. Al and I. D. Yun, “Partial policy-based reinforcement learning for anatomical landmark localization in 3d medical images,” arXiv preprint arXiv:1807.02908, 2018
2018 arXiv
-
[261]
An artificial agent for robust image registration
R. Liao, S. Miao, P. de Tournemire, S. Grbic, A. Kamen, T. Mansi, and D. Comaniciu, “An artificial agent for robust image registration.” in AAAI, 2017, pp. 4168–4175
2017
-
[262]
Multimodal image registration with deep context re- inforcement learning,
K. Ma, J. Wang, V . Singh, B. Tamersoy, Y .-J. Chang, A. Wimmer, and T. Chen, “Multimodal image registration with deep context re- inforcement learning,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2017, pp. 240–248
2017
-
[263]
Robust non-rigid registra- tion through agent-based action learning,
J. Krebs, T. Mansi, H. Delingette, L. Zhang, F. C. Ghesu, S. Miao, A. K. Maier, N. Ayache, R. Liao, and A. Kamen, “Robust non-rigid registra- tion through agent-based action learning,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . ...
2017
-
[264]
Deep reinforcement learning for active breast lesion detection from dce-mri,
G. Maicas, G. Carneiro, A. P. Bradley, J. C. Nascimento, and I. Reid, “Deep reinforcement learning for active breast lesion detection from dce-mri,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2017, pp. 665–673
2017
-
[265]
Deep reinforcement learning for vessel centerline tracing in multi-modality 3d volumes,
P. Zhang, F. Wang, and Y . Zheng, “Deep reinforcement learning for vessel centerline tracing in multi-modality 3d volumes,” in Inter- national Conference on Medical Image Computing and Computer- Assisted Intervention. Springer, 2018, pp. 755–763
2018
-
[266]
Application on reinforcement learning for diag- nosis based on medical image,
S. M. B. Netto, V . R. C. Leite, A. C. Silva, A. C. de Paiva, and A. de Almeida Neto, “Application on reinforcement learning for diag- nosis based on medical image,” in Reinforcement Learning. InTech, 2008
2008
-
[267]
Lead: a methodology for learning efficient approaches to medical diagnosis,
S. J. Fakih and T. K. Das, “Lead: a methodology for learning efficient approaches to medical diagnosis,” IEEE Transactions on Information Technology in Biomedicine, vol. 10, no. 2, pp. 220–228, 2006
2006
-
[268]
Overview of the trec 2016 clinical decision support track
K. Roberts, M. S. Simpson, E. M. V oorhees, and W. R. Hersh, “Overview of the trec 2016 clinical decision support track.” in TREC, 2016
2016
-
[269]
Learning to diagnose: Assimilating clinical narratives using deep reinforcement learning,
Y . Ling, S. A. Hasan, V . Datla, A. Qadir, K. Lee, J. Liu, and O. Farri, “Learning to diagnose: Assimilating clinical narratives using deep reinforcement learning,” in Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Pape...
2017
-
[270]
Breast cancer surveillance consortium: a national mammography screening and outcomes database
R. Ballard-Barbash, S. H. Taplin, B. C. Yankaskas, V . L. Ernster, R. D. Rosenberg, P. A. Carney, W. E. Barlow, B. M. Geller, K. Kerlikowske, B. K. Edwards et al., “Breast cancer surveillance consortium: a national mammography screening and outcomes database.” American Journal...
1997
-
[271]
An adaptive online learning framework for practical breast cancer diagnosis,
T. Chu, J. Wang, and J. Chen, “An adaptive online learning framework for practical breast cancer diagnosis,” in Medical Imaging 2016: Computer-Aided Diagnosis, vol. 9785. International Society for Optics and Photonics, 2016, p. 978524
2016
-
[272]
Inquire and di- agnose: Neural symptom checking ensemble using deep reinforcement learning,
K.-F. Tang, H.-C. Kao, C.-N. Chou, and E. Y . Chang, “Inquire and di- agnose: Neural symptom checking ensemble using deep reinforcement learning,” in Proceedings of NIPS Workshop on Deep Reinforcement Learning, 2016
2016
-
[273]
Context-aware symptom checking for disease diagnosis using hierarchical reinforcement learn- ing,
H.-C. Kao, K.-F. Tang, and E. Y . Chang, “Context-aware symptom checking for disease diagnosis using hierarchical reinforcement learn- ing,” 2018
2018
-
[274]
Artificial intelligence in xprize deepq tricorder,
E. Y . Chang, M.-H. Wu, K.-F. T. Tang, H.-C. Kao, and C.-N. Chou, “Artificial intelligence in xprize deepq tricorder,” in Proceedings of the 2nd International Workshop on Multimedia for Personal Health and Health Care. ACM, 2017, pp. 11–18
2017
-
[275]
Deepq: Advancing healthcare through artificial intel- ligence and virtual reality,
E. Y . Chang, “Deepq: Advancing healthcare through artificial intel- ligence and virtual reality,” in Proceedings of the 2017 ACM on Multimedia Conference. ACM, 2017, pp. 1068–1068
2017
-
[276]
Task-oriented dialogue system for automatic diagnosis,
Z. Wei, Q. Liu, B. Peng, H. Tou, T. Chen, X. Huang, K.-F. Wong, and X. Dai, “Task-oriented dialogue system for automatic diagnosis,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , vol. 2, 2018, pp. 201–207
2018
-
[277]
Improving mild cognitive impairment prediction via reinforcement learning and dialogue simulation,
F. Tang, K. Lin, I. Uchendu, H. H. Dodge, and J. Zhou, “Improving mild cognitive impairment prediction via reinforcement learning and dialogue simulation,” arXiv preprint arXiv:1802.06428 , 2018
2018 arXiv
-
[278]
Approximate dynamic programming for capacity allocation in the service industry,
H.-J. Schuetz and R. Kolisch, “Approximate dynamic programming for capacity allocation in the service industry,” European Journal of Operational Research, vol. 218, no. 1, pp. 239–250, 2012
2012
-
[279]
Reinforcement learning based resource allocation in business process management,
Z. Huang, W. M. van der Aalst, X. Lu, and H. Duan, “Reinforcement learning based resource allocation in business process management,” Data & Knowledge Engineering , vol. 70, no. 1, pp. 127–145, 2011
2011
-
[280]
Clinic scheduling models with overbooking for patients with heterogeneous no-show probabilities,
B. Zeng, A. Turkcan, J. Lin, and M. Lawley, “Clinic scheduling models with overbooking for patients with heterogeneous no-show probabilities,” Annals of Operations Research, vol. 178, no. 1, pp. 121– 144, 2010
2010
-
[281]
Reinforcement learning for primary care e appointment scheduling,
T. S. M. T. Gomes, “Reinforcement learning for primary care e appointment scheduling,” 2017
2017
-
[282]
Asynchronous methods for deep reinforcement learning,
V . Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in ICML, 2016, pp. 1928–1937
2016
-
[283]
A function approximation method for model- based high-dimensional inverse reinforcement learning,
K. Li and J. W. Burdick, “A function approximation method for model- based high-dimensional inverse reinforcement learning,” arXiv preprint arXiv:1708.07738, 2017
2017 arXiv
-
[284]
Multilateral surgical pattern cutting in 2d orthotropic gauze with deep reinforcement learning policies for tensioning,
B. Thananjeyan, A. Garg, S. Krishnan, C. Chen, L. Miller, and K. Goldberg, “Multilateral surgical pattern cutting in 2d orthotropic gauze with deep reinforcement learning policies for tensioning,” in IEEE ICRA’17. IEEE, 2017, pp. 2371–2378
2017
-
[285]
A new tensioning method using deep reinforcement learning for surgical pattern cutting,
T. T. Nguyen, N. D. Nguyen, F. Bello, and S. Nahavandi, “A new tensioning method using deep reinforcement learning for surgical pattern cutting,” arXiv preprint arXiv:1901.03327 , 2019
1901 arXiv
-
[286]
Towards transferring skills to flexible surgical robots with programming by demonstration and reinforcement learning,
J. Chen, H. Y . Lau, W. Xu, and H. Ren, “Towards transferring skills to flexible surgical robots with programming by demonstration and reinforcement learning,” in ICACI’16. IEEE, 2016, pp. 378–384
2016
-
[287]
Path planning for automation of surgery robot based on probabilistic roadmap and reinforcement learning,
D. Baek, M. Hwang, H. Kim, and D.-S. Kwon, “Path planning for automation of surgery robot based on probabilistic roadmap and reinforcement learning,” in 2018 15th International Conference on Ubiquitous Robots (UR) . IEEE, 2018, pp. 342–347
2018
-
[288]
Inverse reinforcement learning via function approximation for clinical motion analysis,
K. Li, M. Rath, and J. W. Burdick, “Inverse reinforcement learning via function approximation for clinical motion analysis,” in IEEE ICRA’18. IEEE, 2018, pp. 610–617
2018
-
[289]
Training an actor-critic reinforcement learn- ing controller for arm movement using human-generated rewards,
K. M. Jagodnik, P. S. Thomas, A. J. van den Bogert, M. S. Bran- icky, and R. F. Kirsch, “Training an actor-critic reinforcement learn- ing controller for arm movement using human-generated rewards,” IEEE Transactions on Neural Systems and Rehabilitation Engineering , vol. 25, ...
1905
-
[290]
Medical qos provision based on reinforcement learning in ultrasound streaming over 3.5 g wireless systems,
R. S. Istepanian, N. Y . Philip, and M. G. Martini, “Medical qos provision based on reinforcement learning in ultrasound streaming over 3.5 g wireless systems,” IEEE Journal on Selected areas in Communications, vol. 27, no. 4, 2009
2009
-
[291]
Cross-layer ultrasound video streaming over mobile wimax and hsupa networks,
A. Alinejad, N. Y . Philip, and R. S. Istepanian, “Cross-layer ultrasound video streaming over mobile wimax and hsupa networks,” IEEE transactions on Information Technology in Biomedicine, vol. 16, no. 1, pp. 31–39, 2012
2012
-
[292]
Functional electrical stimulation after spinal cord injury: current use, therapeutic effects and future directions,
K. Ragnarsson, “Functional electrical stimulation after spinal cord injury: current use, therapeutic effects and future directions,” Spinal cord, vol. 46, no. 4, p. 255, 2008
2008
-
[293]
Creating a reinforcement learning controller for functional electrical stimulation of a human arm,
P. S. Thomas, M. Branicky, A. Van Den Bogert, and K. Jagodnik, “Creating a reinforcement learning controller for functional electrical stimulation of a human arm,” in The Yale Workshop on Adaptive and Learning Systems, vol. 49326. NIH Public Access, 2008, p. 1
2008
-
[294]
Application of the actor-critic architecture to functional electrical stimulation control of a human arm
P. S. Thomas, A. J. van den Bogert, K. M. Jagodnik, and M. S. Branicky, “Application of the actor-critic architecture to functional electrical stimulation control of a human arm.” in IAAI, 2009
2009
-
[295]
Can the pharmaceutical industry reduce attrition rates?
I. Kola and J. Landis, “Can the pharmaceutical industry reduce attrition rates?” Nature reviews Drug discovery , vol. 3, no. 8, p. 711, 2004
2004
-
[296]
Schneider, De novo molecular design
G. Schneider, De novo molecular design . John Wiley & Sons, 2013
2013
-
[297]
Molecular de-novo design through deep reinforcement learning,
M. Olivecrona, T. Blaschke, O. Engkvist, and H. Chen, “Molecular de-novo design through deep reinforcement learning,” Journal of Cheminformatics, vol. 9, no. 1, p. 48, 2017
2017
-
[298]
Accelerating drugs discovery with deep reinforcement learning: An early approach,
A. Serrano, B. Imbern ´on, H. P ´erez-S´anchez, J. M. Cecilia, A. Bueno- Crespo, and J. L. Abell ´an, “Accelerating drugs discovery with deep reinforcement learning: An early approach,” in Proceedings of the 47th International Conference on Parallel Processing Companion . ACM,...
2018
-
[299]
Exploring deep recurrent models with reinforcement learning for molecule design,
D. Neil, M. Segler, L. Guasch, M. Ahmed, D. Plumbley, M. Sellwood, and N. Brown, “Exploring deep recurrent models with reinforcement learning for molecule design,” 2018
2018
-
[300]
Deep reinforcement learning for de novo drug design,
M. Popova, O. Isayev, and A. Tropsha, “Deep reinforcement learning for de novo drug design,” Science Advances, vol. 4, no. 7, p. eaap7885, 2018
2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.