REVIEW 2 major objections 1 minor 71 references
Insulin4RL supplies a dataset of over 375,000 irregular clinical decisions from real ICU insulin trajectories for offline reinforcement learning.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-26 21:16 UTC pith:IEVONC24
load-bearing objection Insulin4RL is a new MIMIC-IV dataset with irregular real-world insulin titration decisions for ORL, but its value hinges on unshown labeling validation. the 2 major comments →
Insulin4RL: Real-Time Insulin Management in the Intensive Care Unit for Offline Reinforcement Learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Insulin4RL is a healthcare offline reinforcement learning dataset featuring naturally irregular inputs and actions from real clinical trajectories, derived from MIMIC-IV, comprising over 375,000 labelled decisions across 12,209 patients requiring insulin infusion titration in the Intensive Care Unit, to enable research into ORL model performance under realistic clinical sampling assumptions.
What carries the argument
The Insulin4RL dataset, which extracts and labels real EHR trajectories while preserving irregular timing of inputs and actions for insulin titration decisions.
Load-bearing premise
The extracted and labeled decisions from MIMIC-IV EHR data accurately represent true clinical decision-making processes and trajectories without substantial labeling errors, missing context, or selection biases introduced during dataset construction.
What would settle it
An independent audit of raw MIMIC-IV notes for a random sample of patients that finds a substantial fraction of the dataset labels do not match documented clinician actions or their exact timings.
If this is right
- Baseline performance metrics using model-free offline reinforcement learning methods are provided for the dataset.
- A standardised evaluation protocol using fitted Q-evaluation is established to assess models.
- The dataset supports direct investigation of ORL performance when clinical sampling occurs at irregular intervals.
- Future research areas identified by the authors can be addressed with this resource.
Where Pith is reading between the lines
- Comparisons between models trained on this irregular dataset and on discretized versions could quantify how much artificial regularity has affected prior results.
- The same extraction approach could be applied to other ICU interventions to create additional irregular-sampling benchmarks.
- If models trained here transfer to live settings, they might reduce reliance on fixed-interval assumptions in deployed clinical decision support.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims to introduce Insulin4RL, a dataset derived from MIMIC-IV containing over 375,000 labelled insulin titration decisions across 12,209 ICU patients. It features naturally irregular sampling from real clinical trajectories (in contrast to temporally discretized EHR datasets), provides a description of structure and characteristics, reports baseline performance using model-free offline RL, and defines a standardized evaluation protocol via fitted Q-evaluation.
Significance. If the labelled decisions are shown to be reliable, the dataset would be a useful resource for ORL research by supporting evaluation under realistic irregular clinical sampling. The inclusion of baselines and a reproducible evaluation protocol is a positive feature that aids future work.
major comments (2)
- [Dataset construction and labeling] The dataset construction section provides no validation (clinician review, inter-rater agreement, or sensitivity analysis) of the heuristics used to map raw MIMIC-IV EHR events to labelled titration decisions and states. This directly affects the central claim that the >375k decisions accurately represent true clinical trajectories without substantial labeling errors or biases from data granularity.
- [Patient cohort and data preprocessing] No details are given on handling of missing data, undocumented verbal orders, or selection biases in the cohort of 12,209 patients; these omissions are load-bearing for claims about realistic clinical sampling assumptions.
minor comments (1)
- The abstract could briefly note the reliance on extraction heuristics as a limitation.
Simulated Author's Rebuttal
We thank the referee for the constructive comments. We address each major comment below and indicate the revisions that will be made to strengthen the manuscript.
read point-by-point responses
-
Referee: [Dataset construction and labeling] The dataset construction section provides no validation (clinician review, inter-rater agreement, or sensitivity analysis) of the heuristics used to map raw MIMIC-IV EHR events to labelled titration decisions and states. This directly affects the central claim that the >375k decisions accurately represent true clinical trajectories without substantial labeling errors or biases from data granularity.
Authors: We agree that external validation of the labeling heuristics would strengthen the central claims. The current manuscript describes the mapping process but does not include clinician review or inter-rater agreement, as these were not feasible given the scale of 375,000 decisions and resource constraints. In revision we will expand the dataset construction section with a more granular description of the heuristics, add a sensitivity analysis on key mapping parameters to quantify potential labeling errors, and include an explicit discussion of biases arising from data granularity. These additions will be incorporated in the revised manuscript. revision: partial
-
Referee: [Patient cohort and data preprocessing] No details are given on handling of missing data, undocumented verbal orders, or selection biases in the cohort of 12,209 patients; these omissions are load-bearing for claims about realistic clinical sampling assumptions.
Authors: We acknowledge that the manuscript lacks sufficient detail on these preprocessing steps. The revised version will expand the methods section to specify how missing data were handled, the approach taken with undocumented verbal orders (including explicit acknowledgment of MIMIC-IV limitations), the exact selection criteria and exclusion rules applied to arrive at the 12,209-patient cohort, and a discussion of potential selection biases. These clarifications will directly support the claims regarding realistic irregular sampling. revision: yes
Circularity Check
No circularity: dataset construction paper with no derivation chain
full rationale
The paper introduces Insulin4RL as a new ORL dataset extracted and labeled from MIMIC-IV EHR records for insulin titration decisions. No mathematical derivations, equations, fitted parameters, or predictions are claimed. The central contribution is the data resource itself (375k labelled decisions across 12k patients), and any labeling heuristics are part of the dataset construction process rather than a result derived from prior steps that reduce to the inputs by construction. This matches the default case of a self-contained data contribution with no load-bearing derivation that could exhibit circularity.
Axiom & Free-Parameter Ledger
read the original abstract
Offline reinforcement learning (ORL) offers the potential to improve the quality of clinical decision-making using historical electronic health record (EHR) data. Current training and evaluative practices in this field rely heavily on EHR datasets that have been temporally discretised into fixed, regular time intervals. Discretisation creates fictional representations of complex clinical scenarios and compromises the generalisability of retrospective model evaluations. In this paper, we introduce Insulin4RL, a healthcare ORL dataset featuring naturally irregular inputs and actions from real clinical trajectories. Derived from MIMIC-IV, Insulin4RL comprises over 375,000 labelled decisions across 12,209 patients requiring insulin infusion titration in the Intensive Care Unit. The dataset can thus be used for research into ORL model performance under realistic clinical sampling assumptions. We provide a description of the dataset's structure and characteristics, baseline performance metrics using model-free offline reinforcement learning, and a standardised evaluation protocol using fitted Q-evaluation. We conclude with suggested areas for future research that could be addressed using this resource.
Figures
Reference graph
Works this paper leans on
-
[1]
Nature Medicine , year =
Komorowski, Matthieu and Celi, Leo A and Badawi, Omar and Gordon, Anthony C and Faisal, A Aldo , title =. Nature Medicine , year =
-
[2]
2022 , journal =
Kumar, Aviral and Hong, Joey and Singh, Anikait and Levine, Sergey , title =. 2022 , journal =
2022
-
[3]
Discretizing Logged Interaction Data Biases Learning for Decision-Making
Schulam, Peter and Saria, Suchi , title =. arXiv preprint arXiv:1810.03025 , year =. doi:10.48550/arXiv.1810.03025 , url =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.1810.03025
-
[4]
Jeter, Russell and Josef, Christopher and Shashikumar, Supreeth and Nemati, Shamim , title =. arXiv preprint arXiv:1902.03271 , year =. doi:10.48550/arXiv.1902.03271 , url =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.1902.03271 1902
-
[5]
Is Deep Reinforcement Learning Ready for Practical Applications in Healthcare?
Lu, Mingyu and Shahn, Zachary and Sow, Daby and Doshi. Is Deep Reinforcement Learning Ready for Practical Applications in Healthcare?. 2020 , month = nov, volume =
2020
-
[6]
npj Digital Medicine , year =
Wu, XiaoDan and Li, RuiChang and He, Zhen and Yu, TianZhi and Cheng, ChangQing , title =. npj Digital Medicine , year =
-
[7]
RLC 2025 Workshop on Practical Insights into Reinforcement Learning for Real Systems , year =
Sun, Yingchuan and Tang, Shengpu , title =. RLC 2025 Workshop on Practical Insights into Reinforcement Learning for Real Systems , year =
2025
-
[8]
Making deep
Tallec, Corentin and Blier, L. Making deep. Proceedings of the 36th. 2019 , month = jun, volume =
2019
-
[9]
and Hajaj, Nissan and Hardt, Michaela and Liu, Peter J
Rajkomar, Alvin and Oren, Eyal and Chen, Kai and Dai, Andrew M. and Hajaj, Nissan and Hardt, Michaela and Liu, Peter J. and Liu, Xiaobing and Marcus, Jake and Sun, Mimi and Sundberg, Patrik and Yee, Hector and Zhang, Kun and Zhang, Yi and Flores, Gerardo and Duggan, Gavin E. and Irvine, Jamie and Le, Quoc and Litsch, Kurt and Mossin, Alexander and Tansuwa...
-
[10]
Proceedings of the 2014 conference on empirical methods in natural language processing (
Kim, Yoon , title =. Proceedings of the 2014 conference on empirical methods in natural language processing (. 2014 , month = oct, pages =
2014
-
[11]
2020 , month = dec, volume =
Kumar, Aviral and Zhou, Aurick and Tucker, George and Levine, Sergey , title =. 2020 , month = dec, volume =
2020
-
[12]
Identifying Decision Points for Safe and Interpretable Reinforcement Learning in Hypotension Treatment , booktitle =
Zhang, Kristine and Wang, Yuanheng and Du, Jianzhun and Chu, Brian and Celi, Leo Anthony and Kindle, Ryan and Doshi. Identifying Decision Points for Safe and Interpretable Reinforcement Learning in Hypotension Treatment , booktitle =. 2021 , month = dec, volume =
2021
-
[13]
Proceedings of the 1st Machine Learning in Health Care , series =
Lipton, Zachary C and Kale, David C and Wetzel, Randall and others , title =. Proceedings of the 1st Machine Learning in Health Care , series =. 2016 , month = aug, volume =
2016
-
[14]
and Hong, Zhang-Wei and Sabounchi, Moein and Sawant, Ashwin S
Desman, Jacob M. and Hong, Zhang-Wei and Sabounchi, Moein and Sawant, Ashwin S. and Gill, Jaskirat and Costa, Ana C. and Kumar, Gagan and Sharma, Rajeev and Gupta, Arpeta and McCarthy, Paul and Nandwani, Veena and Powell, Doug and Carideo, Alexandra and Goodwin, Donnie and Ahmed, Sanam and Gidwani, Umesh and Levin, Matthew A. and Varghese, Robin and Filso...
-
[15]
Proceedings of the 33rd
Jiang, Nan and Li, Lihong , title =. Proceedings of the 33rd. 2016 , month = jul, volume =
2016
-
[16]
Proceedings of the 36th
Le, Hoang and Voloshin, Cameron and Yue, Yisong , title =. Proceedings of the 36th. 2019 , month = jun, volume =
2019
-
[17]
The Health Gym: synthetic health-related datasets for the development of reinforcement learning algorithms , journal =
Kuo, Nicholas I-Hsien and Polizzotto, Mark N and Finfer, Simon and Garcia, Federico and S. The Health Gym: synthetic health-related datasets for the development of reinforcement learning algorithms , journal =. 2022 , month = nov, volume =
2022
-
[18]
and Bellomo, Rinaldo and Cousins, Caroline E
Plummer, Mark P. and Bellomo, Rinaldo and Cousins, Caroline E. and Annink, Christopher E. and Sundararajan, Krishnaswamy and Reddi, Benjamin A. J. and Raj, John P. and Chapman, Marianne J. and Horowitz, Michael and Deane, Adam M. , title =. Intensive care medicine , year =
-
[19]
and Ketcham, Scott W
Adie, Sarah K. and Ketcham, Scott W. and Marshall, Vincent D. and Farina, Nicholas and Sukul, Devraj , title =. Journal of Diabetes and its Complications , year =
-
[20]
Desgrouas, Maxime and Demiselle, Julien and Stiel, Laure and Brunot, Vincent and Marnai, Rémy and Sarfati, Sacha and Fiancette, Maud and Lambiotte, Fabien and Thille, Arnaud W. and Leloup, Maxime and Clerc, Sébastien and Beuret, Pascal and Bourion, Anne-Astrid and Daum, Johan and Malhomme, Rémi and Ravan, Ramin and Sauneuf, Bertrand and Rigaud, Jean-Phili...
-
[21]
Data-driven curation process for describing the blood glucose management in the intensive care unit , journal =
Robles Ar. Data-driven curation process for describing the blood glucose management in the intensive care unit , journal =. 2021 , month = mar, volume =
2021
-
[22]
2022 , month = apr, doi =
Kostrikov, Ilya and Nair, Ashvin and Levine, Sergey , title =. 2022 , month = apr, doi =
2022
-
[23]
2023 , month = oct, volume =
Zhu, Taiyu and Li, Kezhi and Georgiou, Pantelis , title =. 2023 , month = oct, volume =
2023
-
[24]
2009 , month = mar, volume =
Intensive versus conventional glucose control in critically ill patients , journal =. 2009 , month = mar, volume =
2009
-
[25]
Tight blood-glucose control without early parenteral nutrition in the ICU , journal =
Gunst, Jan and Debaveye, Yves and G\". Tight blood-glucose control without early parenteral nutrition in the ICU , journal =. 2023 , month = sep, volume =
2023
-
[26]
Scientific Reports , year =
Che, Zhengping and Purushotham, Sanjay and Cho, Kyunghyun and Sontag, David and Liu, Yan , title =. Scientific Reports , year =
-
[27]
2020 , month = dec, volume =
Kidger, Patrick and Morrill, James and Foster, James and Lyons, Terry , title =. 2020 , month = dec, volume =
2020
-
[28]
2021 , month = may, doi =
Shukla, Satya Narayan and Marlin, Benjamin M , title =. 2021 , month = may, doi =
2021
-
[29]
, title =
Tipirneni, Sindhu and Reddy, Chandan K. , title =. ACM Transactions on Knowledge Discovery from Data , year =
-
[30]
Neural Computation , year =
Hochreiter, Sepp and Schmidhuber, J. Neural Computation , year =
-
[31]
and Bossi, Luca and Dalla Man, Chiara and De Nicolao, Giuseppe and Kovatchev, Boris and Cobelli, Claudio , title =
Magni, Lalo and Raimondo, Davide M. and Bossi, Luca and Dalla Man, Chiara and De Nicolao, Giuseppe and Kovatchev, Boris and Cobelli, Claudio , title =. Journal of Diabetes Science and Technology , year =
-
[32]
Indiana University Mathematics Journal , year =
Bellman, Richard , title =. Indiana University Mathematics Journal , year =
-
[33]
Medical event data standard (
Arnrich, Bert and Choi, Edward and Fries, Jason Alan and McDermott, Matthew BA and Oh, Jungwoo and Pollard, Tom and Shah, Nigam and Steinberg, Ethan and Wornow, Michael and. Medical event data standard (. 2024 , month = mar, url =
2024
-
[34]
and Pollard, Tom J
Johnson, Alistair E.W. and Pollard, Tom J. and Shen, Lu and Lehman, Li-wei H. and Feng, Mengling and Ghassemi, Mohammad and Moody, Benjamin and Szolovits, Peter and Celi, Leo A. and Mark, Roger G. , title =. Scientific Data , year =
-
[35]
Machine learning for healthcare conference , year =
Raghu, Aniruddh and Komorowski, Matthieu and Celi, Leo Anthony and Szolovits, Peter and Ghassemi, Marzyeh , title =. Machine learning for healthcare conference , year =
-
[36]
Johnson, Alistair E. W. and Bulgarelli, Lucas and Shen, Lu and Gayles, Alvin and Shammout, Ayad and Horng, Steven and Pollard, Tom J. and Hao, Sicheng and Moody, Benjamin and Gow, Brian and Lehman, Li-wei H. and Celi, Leo A. and Mark, Roger G. , title =. Scientific Data , year =
-
[37]
MIMIC-IV.PhysioNet, October 2024
Johnson, Alistair and Bulgarelli, Lucas and Pollard, Tom and Gow, Brian and Moody, Benjamin and Horng, Steven and Celi, Leo Anthony and Mark, Roger , title =. 2024 , howpublished =. doi:10.13026/kpb9-mt58 , note =
-
[38]
2025 IEEE EMBS International Conference on Biomedical and Health Informatics (BHI) , year =
Huang, Yong and Yang, Zhongqi and Rahmani, Amir , title =. 2025 IEEE EMBS International Conference on Biomedical and Health Informatics (BHI) , year =
2025
-
[39]
Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks , year =
Voloshin, Cameron and Le, Hoang Minh and Jiang, Nan and Yue, Yisong , title =. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks , year =
-
[40]
Proceedings of the 6th Machine Learning for Healthcare Conference , series =
Tang, Shengpu and Wiens, Jenna , title =. Proceedings of the 6th Machine Learning for Healthcare Conference , series =. 2021 , month = aug, volume =
2021
-
[41]
and Komorowski, Matthieu and Komorowski, Matthieu and Faisal, Aldo and Celi, Leo Anthony and Sontag, David and Doshi-Velez, Finale , title =
Gottesman, Omer and Johansson, Fredrik and Meier, Joshua and Dent, Jack and Lee, Donghun and Srinivasan, Srivatsan and Zhang, Linying and Ding, Yi and Wihl, David and Peng, Xuefeng and Yao, Jiayu and Lage, Isaac and Mosch, Christopher and Lehman, Li-wei H. and Komorowski, Matthieu and Komorowski, Matthieu and Faisal, Aldo and Celi, Leo Anthony and Sontag,...
2018
-
[42]
CEUR Workshop Proceedings , year =
Marling, Cindy and Bunescu, Razvan , title =. CEUR Workshop Proceedings , year =
-
[43]
2020 , journal =
Levine, Sergey and Kumar, Aviral and Tucker, George and Fu, Justin , title =. 2020 , journal =
2020
-
[44]
European conference on machine learning , year =
Frank, Eibe and Hall, Mark , title =. European conference on machine learning , year =
-
[45]
and Stone, Peter , title =
Hausknecht, Matthew J. and Stone, Peter , title =. 2015 , month = nov, pages =
2015
-
[46]
Proceedings of the 39th
Ni, Tianwei and Eysenbach, Benjamin and Salakhutdinov, Ruslan , title =. Proceedings of the 39th. 2022 , month = jul, volume =
2022
-
[47]
Proceedings of the 35th
Igl, Maximilian and Zintgraf, Luisa and Le, Tuan Anh and Wood, Frank and Whiteson, Shimon , title =. Proceedings of the 35th. 2018 , month = jul, volume =
2018
-
[48]
Operations Research , year =
White, Chelsea C , title =. Operations Research , year =
-
[49]
Yu, Huizhen , title =
-
[50]
IEEE Transactions on Reliability , year =
Zhang, Mimi and Revie, Matthew , title =. IEEE Transactions on Reliability , year =
-
[51]
Individualised versus conventional glucose control in critically-ill patients: the
Boh. Individualised versus conventional glucose control in critically-ill patients: the. Intensive Care Medicine , year =
-
[52]
and Hermanides, Jeroen and Deane, Adam M
Plummer, Mark P. and Hermanides, Jeroen and Deane, Adam M. , title =. Current Opinion in Clinical Nutrition & Metabolic Care , year =
-
[53]
Circulation , year =
Goldberger, Ary L and Amaral, Luis AN and Glass, Leon and Hausdorff, Jeffrey M and Ivanov, Plamen Ch and Mark, Roger G and Mietus, Joseph E and Moody, George B and Peng, Chung-Kang and Stanley, H Eugene , title =. Circulation , year =
-
[54]
Journal of Medical Internet Research , year =
Liu, Siqi and See, Kay Choong and Ngiam, Kee Yuan and Celi, Leo Anthony and Sun, Xingzhi and Feng, Mengling , title =. Journal of Medical Internet Research , year =
-
[55]
npj Digital Medicine , year =
Jayaraman, Pushkala and Desman, Jacob and Sabounchi, Moein and Nadkarni, Girish N and Sakhuja, Ankit , title =. npj Digital Medicine , year =
-
[56]
Nature Medicine , year =
Wang, Guangyu and Liu, Xiaohong and Ying, Zhen and Yang, Guoxing and Chen, Zhiwei and Liu, Zhiwen and Zhang, Min and Yan, Hongmei and Lu, Yuxing and Gao, Yuanxu and Xue, Kanmin and Li, Xiaoying and Chen, Ying , title =. Nature Medicine , year =
-
[57]
and Chowdhury, Afsana and Feng, Guangyu and Peters, Jennifer J
Gao, Qitong and Schmidt, Stephen L. and Chowdhury, Afsana and Feng, Guangyu and Peters, Jennifer J. and Genty, Katherine and Grill, Warren M. and Turner, Dennis A. and Pajic, Miroslav , title =. International Conference on Cyber-Physical Systems , year =
-
[58]
2026 , month = feb, volume =
Fan, Fan and Huang, Hao and Yan, Jingwen and Xu, Chao-Yue and Wu, Xiuhua and Zhou, Chunmei and Wen, Dandan and Huang, Hai and Li, Ho Cheung and Qiu, Yihong , title =. 2026 , month = feb, volume =
2026
-
[59]
and Wiens, Jenna , title =
Tang, Shengpu and Modi, Aditya and Sjoding, Michael W. and Wiens, Jenna , title =. Proceedings of the 37th. 2020 , month = jul, volume =
2020
-
[60]
Transatlantic transferability of a new reinforcement learning model for optimizing haemodynamic treatment for critically ill patients with sepsis , journal =
Roggeveen, Luca and. Transatlantic transferability of a new reinforcement learning model for optimizing haemodynamic treatment for critically ill patients with sepsis , journal =. 2021 , month = feb, volume =
2021
-
[61]
and Subramanian, Jayakumar and Ghassemi, Marzyeh , title =
Fatemi, Mehdi and Killian, Taylor W. and Subramanian, Jayakumar and Ghassemi, Marzyeh , title =. 2021 , month = dec, volume =
2021
-
[62]
Human-Centric Intelligent Systems , year =
Tu, Rui and Luo, Zhipeng and Pan, Chuanliang and Wang, Zhong and Su, Jie and Zhang, Yu and Wang, Yifan , title =. Human-Centric Intelligent Systems , year =
-
[63]
Operations Research , year =
Jewell, William S , title =. Operations Research , year =
-
[64]
, title =
Howard, Ronald A. , title =. Bulletin of the International Statistical Institute , year =
-
[65]
Advances in Neural Information Processing Systems , year =
Bradtke, Steven and Duff, Michael , title =. Advances in Neural Information Processing Systems , year =
-
[66]
Proceedings of the 41st
Farebrother, Jesse and Orbay, Jordi and Vuong, Quan and Ali Taiga, Adrien and Chebotar, Yevgen and Xiao, Ted and Irpan, Alex and Levine, Sergey and Castro, Pablo Samuel and Faust, Aleksandra and Kumar, Aviral and Agarwal, Rishabh , title =. Proceedings of the 41st. 2024 , month = jul, volume =
2024
-
[67]
Sutton, Richard S and Barto, Andrew G , title =
-
[68]
2026 , journal =
Frost, Thomas and Vaidya, Hrisheekesh and Harris, Steve , title =. 2026 , journal =
2026
-
[69]
Journal of Open Source Software , year =
McInnes, Leland and Healy, John and Saul, Nathaniel and Großberger, Lukas , title =. Journal of Open Source Software , year =
-
[70]
Intensive insulin therapy in critical care: a review of 12 protocols , journal =
Wilson, Mark and Weinreb, Jane and. Intensive insulin therapy in critical care: a review of 12 protocols , journal =. 2007 , month = apr, volume =
2007
-
[71]
and Singh, Satinder P
Precup, Doina and Sutton, Richard S. and Singh, Satinder P. , title =. Proceedings of the 17th. 2000 , pages =
2000
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.