{"id":"1e30cdd3-683b-4157-a725-cc7954a71299","arxiv_id":"2411.16346","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper introduces a harmonized multi-center critical care time series dataset and transfer benchmark covering nine datasets from three continents, with treatment variables, and compares seven models on early event prediction tasks.","lead":"A team from ETH Zurich and partner hospitals combined nine public critical care datasets from the US, Europe, and China into one harmonized time series benchmark that also includes treatment variables. The paper shows gradient-boosted trees still beat deep sequence models on most tasks, while pretraining on all datasets helps transfer to new hospitals with limited data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing premise is that ricu-based treatment concepts (Sect. 3.2, Tables 6-8) preserve clinical meaning across sites: a harmonized norepinephrine rate must denote the same infusion in MIMIC-IV, eICU, HiRID, SICdb, and PICdb.","rationale":"The reader's weakest assumption is also the most load-bearing condition I can find, so my analysis agrees rather than adding a new objection. For the central claim to be true, the 141-concept schema of Table 8, especially the 55 treatment variables with administration rates, must denote the same clinical quantity in every source dataset. Section 3.2 describes the harmonization as expert opinion informed by literature importance, with no quantitative validation, and the paper's own limitation statement (Appendix D) does not flag this gap. The stakes are high because the transfer benchmark in Section 4.3 is the experimental demonstration that harmonization enables generalization; if treatment concepts are site-incomparable, the Table 4 AUROC values could reflect charting intensity, unit conventions, or the label-prevalence differences visible in Table 5 rather than shared physiology. I nevertheless credit the paper's genuine strengths: a materially larger multi-center collection, a broad and expensive benchmark (about 5124 runs across seven architectures), a useful supervised fine-tuning study (Figures 3 and 5), an explicit limitations section, and honest acknowledgment of BlendedICU (Oliver et al. 2023). The internal inconsistencies the reader noted (Table 7 columns that do not sum; Table 4 SICdb LR multi-center values of 95.7 that exceed both the single-center and the strongest model values) are real and support the call for corrected tables, but they do not by themselves prove the harmonization is wrong. Because the key concern is falsifiable by an independent recomputation that uses only the source data and the published concept table, and because the paper can be repaired by releasing the pipeline and auditing the mappings, the appropriate verdict is the reader's CONDITIONAL one, unchanged.","tokens_in":27531,"tokens_out":16367,"duration_ms":141542,"concrete_test":"Independently validate the norepinephrine rate concept (Table 8, mcg/min). For a stratified random sample of roughly 50 stays per dataset, recompute the rate directly from raw source administration tables (MIMIC-IV inputevents, eICU infusionDrug, HiRID drug-order events, and the SICdb and PICdb medication records), applying the documented units and weight conversions, and compare against the harmonized time grid at matched timestamps. Accept the mapping only if recomputed and harmonized values agree within ±10% at at least 95% of timepoints in every dataset; a secondary check compares weight-normalized dose distributions among hypotensive mechanically ventilated adult patients, requiring adult interquartile ranges to overlap within a factor of two. Failure of either criterion means the harmonized treatment claim and the transfer benchmark are confounded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (abstract, Section 1) is a dataset contribution: the largest harmonized critical care time-series collection, the first to harmonize core treatment variables, to combine ICU and ED, and to include Asia. For that claim to hold, the expert-defined ricu concepts of Section 3.2 (Tables 6 and 8) must make a norepinephrine rate, propofol rate, or heparin rate clinically comparable across MIMIC-IV, eICU, HiRID, SICdb, UMCdb, PICdb, and Zigong. Constructing these rates requires correct drug identification, unit conversion (including weight-dependent mcg/kg/min to mcg/min, delicate for pediatric PICdb), bolus-versus-infusion disambiguation, and time alignment; the paper offers no audit of these steps, no dual expert coding, no reference-standard comparison, and no dose-response sanity check. If rates are mis-scaled or mis-mapped, the 'harmonized core treatment variables' contribution fails on its own terms, and the out-of-distribution transfer results (Section 4.3, Table 4) no longer measure clinical generalization: they could reflect site-specific charting artifacts, label prevalence differences (Table 5 shows positive-label rates from 0% to 54%), or patient-mix differences rather than shared physiology. Internal signals of under-auditing: Table 7's Used/Not used/Total columns do not reconcile for any dataset (e.g., MIMIC-IV: 442/292/453, where 442+292 is not 453), and Table 6 lists vasopressin as High clinical importance with no supporting task or literature citation. The paper itself (Section 3.2) concedes Oliver et al. (2023) already harmonized treatment indicators, so the defensible novelty is narrower (rates plus abstract grouping), which raises the bar for validating those rates. The dataset may still be a useful concatenated resource, but the signature claim is unverified as posted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript introduces a harmonized multi-center critical care time-series dataset assembled from nine publicly available ICU and ED sources (MIMIC-III/IV, eICU, UMCdb, HiRID, SICdb, PICdb, Zigong EHR, and MIMIC-IV-ED), totaling roughly 600,000 stays, and reports an extensive benchmark of seven machine learning models on four early-event-prediction tasks. The authors claim the dataset is the largest harmonized critical care time-series collection, the first to harmonize core treatment variables as both indicators and administration rates, the first to combine ICU and ED data, and the first to include data from Asia. They evaluate in-distribution performance, leave-one-dataset-out transfer, and a supervised pretraining/fine-tuning study on HiRID.","tokens_in":27887,"tokens_out":5325,"duration_ms":53564,"significance":"If the claims hold, this would be a valuable community resource: a multi-continent, multi-unit harmonized dataset with treatment variables, plus a transfer benchmark far broader than prior single-center or smaller multi-center efforts. The experimental design is genuinely extensive, with seven models, four tasks, nine datasets, both single- and multi-center training, and a fine-tuning study using three seeds. The observed result that gradient-boosted trees with feature extraction remain competitive with modern deep sequence models, and that supervised pretraining helps for small target cohorts, is useful and credible. However, the central dataset claim is currently unverifiable because no dataset or code is released, and the treatment harmonization that underpins the headline contribution is not audited. The paper is a promising benchmark report, but it is not yet a verifiable dataset contribution.","major_comments":[{"comment":"The paper is framed as a dataset contribution, yet it provides no dataset access mechanism, URL, repository, or code release anywhere in the manuscript. Without the harmonized artifacts or a clear procedure for obtaining them, the central claims (largest harmonized dataset, first to harmonize treatment variables, first to include Asia) cannot be checked and the benchmark cannot be reproduced. This is load-bearing for the paper's main contribution and must be addressed, either by releasing the processed data and code or by explicitly stating access conditions and providing a public placeholder/release plan.","section":"Abstract and Section 3.3"},{"comment":"The load-bearing premise of the treatment harmonization is that ricu-based concepts such as norepinephrine rate, propofol rate, and heparin rate are clinically comparable across MIMIC-IV, eICU, HiRID, UMCdb, SICdb, PICdb, and Zigong. The paper gives no audit of drug identification, unit conversion (including weight-based mcg/kg/min to mcg/min conversion in pediatric PICdb), bolus-versus-infusion disambiguation, or time alignment, and no external validation such as dual expert coding, reference-standard comparison, or dose-response sanity checks. If rates are mis-scaled or mis-mapped, the out-of-distribution transfer results in Section 4.3 no longer measure clinical generalization but instead reflect site-specific charting artifacts. The paper needs either a detailed validation appendix for the treatment concepts or a clear statement of which treatment concepts are validated and how.","section":"Section 3.2, Tables 6-8"},{"comment":"The Used/Not used/Total columns do not reconcile for any dataset. For example, MIMIC-IV reports 442 used, 292 not used, and 453 total, but 442 + 292 = 734, not 453; the same inconsistency appears in every row, including the Total row (26,143 + 20,943 = 47,086, not 26,264). This table is central evidence for the treatment harmonization claim, and as presented it is internally inconsistent. The authors must clarify what the counts represent and correct the arithmetic.","section":"Table 7"},{"comment":"Several in-distribution multi-center values for SICdb are implausible. For Circulatory 8h, multi-center logistic regression reports 95.7 AUROC while multi-center LightGBM with features reports 91.7 and the best single-center model reports 91.6; for Kidney 48h, multi-center logistic regression reports 93.4 while all other multi-center models are around 89-90. These look like transposed or copied values. Because Tables 2 and 4 are the main benchmark evidence, the authors need to audit all numbers and provide a corrected table, including standard deviations where the text promises them.","section":"Table 2"},{"comment":"Table 5 lists a 0% label prevalence for Zigong on Respiratory 24h and Kidney 48h, yet the paper states in Section 1 that it provides annotations and results on multiple organ failure tasks on the same data. If no positive labels exist for these tasks on Zigong, the absence of Zigong respiratory and kidney rows in Tables 2, 4, 9, and 10 should be explicitly explained; if the 0% entries are errors, they must be corrected. As written, the task coverage claim and the label statistics are in tension.","section":"Table 5 and Section 4"}],"minor_comments":[{"comment":"There are several typographical errors in the treatment concept tables (e.g., 'Antibotics' instead of 'Antibiotics', 'Elektrolytes-Mg' instead of 'Electrolytes-Mg', 'teophyllin' instead of 'theophylline') that should be corrected for a dataset reference document.","section":"Table 6"},{"comment":"The concept reference table lists the indicator concepts dobu_ind, levo_ind, norepi_ind, epi_ind, milrin_ind, teophyllin_ind, dopa_ind, adh_ind, hep_ind, prop_ind, benzdia_ind, and loop_diur_ind twice, which makes the table hard to use as a reference.","section":"Table 8"},{"comment":"Each panel of Figure 3 appears to contain two y-axes with different scales (one around 0.89-0.93 and another around 0.5-0.8), but the axes are not labeled separately in the caption; please clarify which curve uses which scale.","section":"Figure 3"},{"comment":"The sentence claiming to include 'all ICU datasets that are freely available to the academic community' is immediately qualified by the exclusion of RICD because it is not free access; please rephrase to 'all datasets that are freely available without additional contracts' to avoid the apparent contradiction.","section":"Section 3.2"},{"comment":"The mean length of stay for MIMIC-IV is reported as 11 days, which is much higher than typical reported values for MIMIC-IV ICU stays (usually several days); please verify this number and the corresponding mean LoS values for other datasets.","section":"Table 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a workshop-style preprint whose main contribution is a dataset and benchmark. The most serious issue is that the dataset and code are not released, so the central claim cannot be verified; if the authors cannot release the data, they should reframe the paper as a benchmark description rather than a dataset contribution. The Table 7 arithmetic and the suspicious Table 2 SICdb values also need correction. I found no evidence of bad faith, and the benchmark scope and fine-tuning study are genuinely useful."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a substantial dataset-harmonization and benchmark effort with one overclaim in the abstract and one load-bearing assumption that is plausibly fine but unvalidated. The new content is real: the authors fold SICdb, PICdb, Zigong EHR, and MIMIC-IV-ED into the ricu framework, pushing the harmonized collection to ~600k stays across three continents, and they add treatment rates on top of the indicator-level harmonization that Oliver et al. already did. The benchmark is extensive — seven models, four tasks, in-distribution and leave-one-dataset-out transfer, plus a fine-tuning study — and the transfer heatmaps are a useful public artifact even if the headline numbers are not yet independently checkable.\n\nThe main overclaim is in the abstract: 'first large-scale collection to include core treatment variables.' That contradicts the paper's own citation of Oliver et al. (2023), who harmonized treatment indicators. The defensible novelty is the rates and abstract grouping, not treatment variables per se. Easy fix, but it should be made.\n\nThe load-bearing assumption, which is also the main scientific risk, is that a harmonized norepinephrine rate means the same clinical intervention in MIMIC-IV, eICU, HiRID, SICdb, and PICdb. The paper gives no audit of unit conversions (the pediatric weight-based conversion in PICdb is delicate), no bolus-versus-infusion disambiguation check, and no dose-response sanity test. If those steps are wrong, the transfer benchmark measures site-specific charting artifacts, not clinical generalization. I don't see evidence they are wrong — the design is sensible — but the paper does not yet establish them.\n\nTwo smaller credibility flags: Table 7's Used/Not used/Total columns do not reconcile for any dataset (MIMIC-IV: 442/292/453), and Table 6 lists vasopressin as High clinical importance with no supporting literature citation. Both are fixable without new experiments.\n\nThe biggest practical weakness is that the dataset and code are not released, so the core artifact is not checkable. That is serious for a dataset paper, though this is a workshop version and the authors say they plan to release.\n\nBottom line: this deserves a serious referee. A good reviewer would ask for the harmonization pipeline, a validation study of the treatment rates, corrected tables, a softened abstract, and — critically — the actual data release. If those land, this becomes a genuinely useful community resource. As posted, it's a solid reference point for anyone building on ricu, but not something I would build on or cite yet.","headline":"Solid harmonization-and-benchmark work with an overclaim and an unvalidated core assumption; worth reviewing, not worth citing yet.","tokens_in":19,"tokens_out":3167,"would_cite":false,"duration_ms":69568,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces the largest harmonized critical care time series dataset to date—about 600,000 stays from nine ICU and emergency department databases on three continents—and benchmarks how well machine learning models transfer…","keywords":["critical care time series","dataset harmonization","treatment variables","transfer learning","early event prediction","foundation models","multi-center ICU data","electronic health records"],"falsifier":"Take a core treatment concept such as norepinephrine rate and compare its distribution, its relation to patient physiology, and its association with outcomes across MIMIC-IV, HiRID, SICdb, and UMCdb; if the harmonized rates show systematically different dosing conventions or the same code maps to different actual drugs at different sites, the clinical comparability behind the transfer numbers breaks down.","tokens_in":27344,"feed_emoji":"🏥","tokens_out":12552,"duration_ms":101427,"temperature":0.7,"pith_summary":"The paper aims to lay the groundwork for foundation models on critical care time series by assembling a single harmonized dataset from nine publicly available ICU and emergency department databases, spanning the USA, Europe, and China and containing roughly 600,000 patient stays and close to one billion extracted data points. It claims this is the first such collection to harmonize core treatment variables rather than only vital signs and lab values, to include both ICU and ED data, and to incorporate Asian data alongside North American and European data. On top of the dataset, the authors build a transfer learning benchmark for early event prediction of circulatory, respiratory, and kidney failure and decompensation, covering in-distribution, held-out-hospital, and fine-tuning settings. A sympathetic reading is that, if the harmonization preserves clinical meaning across sites, the collection gives the field a reusable resource for studying distribution shift and for pretraining models that smaller hospitals can fine-tune.","feed_headline":"600,000 critical care stays harmonized across three continents","feed_subtitle":"First ICU/ED collection with harmonized treatments, plus a transfer benchmark for early event prediction.","key_machinery":"The central object is the harmonized concept layer built on the ricu package's data-source-agnostic concept abstraction. A concept is one clinical quantity—heart rate, norepinephrine rate, antibiotic indicator, and so on—that has a source-variable mapping for each underlying database. The paper's extension defines 141 concepts: 6 static demographics, 80 observations, and 55 treatment variables, grouping individual drugs into wider treatment concepts and adding both indicators and, for core medications, administration rates. This layer does the load-bearing work: it turns nine heterogeneous hospital schemas into a common grid, lets models train jointly across sites, and is what makes the transfer benchmark interpretable as transfer of clinical meaning rather than mere format.","core_discovery":"The central claim, on the paper's own terms, is that a sufficiently broad and carefully harmonized multi-center dataset makes cross-hospital transfer for critical care time series viable. The authors construct that dataset by mapping local coding schemes into shared concepts using the ricu abstraction layer, extending it with new observation concepts, treatment variables, and the SICdb, PICdb, Zigong EHR, and MIMIC-IV-ED sources. Treatment harmonization is the novel piece: individual drugs are grouped into abstract concepts such as vasopressors, sedatives, antibiotics, and diuretics, with administration rates recorded for core medications like norepinephrine, epinephrine, and propofol. The benchmark then shows that gradient-boosted trees with hand-extracted features are the strongest predictors, that deep sequence models come within one or two AUROC points, and that supervised pretraining on the other hospitals improves performance on a new hospital for small and medium training-set sizes. The authors do not claim a finished foundation model; they claim the foundation: the dataset, task annotations, and transfer results needed to build one.","pith_inferences":["Editorial extension: the treatment harmonization could be externally validated by comparing norepinephrine rate distributions and their association with blood-pressure changes across hospitals; that would test whether 'norepinephrine' means the same intervention everywhere.","Editorial extension: a controlled ablation on the harmonized dataset, training with and without the 55 treatment concepts, would isolate whether treatments help transfer or simply add site-specific confounds, something the current benchmark does not separate.","Editorial extension: the tokenized data format opens a natural next experiment in self-supervised pretraining, such as masked event prediction, before fine-tuning on the early-event tasks; the paper provides the data and benchmark but does not run that pretraining."],"forward_implications":["Training on the harmonized collection should become the reference starting point for small- and medium-scale critical care time series studies, since supervised pretraining on the other hospitals outperformed training from scratch on HiRID up to tens of thousands of admissions.","The transfer heatmaps supply a quantitative map of which hospital pairs are clinically similar: models transfer well within US/European groups and poorly to Chinese and pediatric datasets, so dataset selection for a new site can be guided by recorded resolution and geography.","The treatment concepts, with indicators plus rates for core vasopressors, sedatives, antibiotics, and diuretics, make it possible to study whether and how medication variables affect predictive accuracy, addressing the earlier finding that including them can hurt performance.","Because task labels use clinically defined early-event horizons of 8 hours for circulatory failure, 24 hours for decompensation and respiratory failure, and 48 hours for kidney failure, the benchmark gives future foundation-model work standardized endpoints comparable to prior single-center benchmarks.","Joint ED-ICU modeling becomes feasible for the first time on harmonized data, potentially leading to a unified prediction model that works regardless of hospital unit."],"supporting_citations":[{"why":"Supplies the ricu data-source-agnostic concept abstraction that maps local variables to shared clinical concepts; the harmonization layer is built on it.","marker":"Bennett et al. [2023]"},{"why":"Prior multi-center ICU benchmark whose roughly 400,000 stays and concept set this work expands with new datasets and treatments.","marker":"Van De Water et al. [2023]"},{"why":"Prior international transfer study for sepsis that defines the state of the art this benchmark extends.","marker":"Moor et al. [2023b]"},{"why":"Provides the hand-extracted feature set that the paper extends; these features power the best-performing LightGBM baseline.","marker":"Soenksen et al. [2022]"},{"why":"Defines the circulatory failure early-event prediction task, its 8-hour horizon, and model baselines used in the benchmark.","marker":"Hyland et al. [2020]"},{"why":"Provides MIMIC-IV, one of the principal US ICU data sources in the harmonized collection.","marker":"Johnson et al. [2023b]"},{"why":"Provides eICU, the large US multi-center ICU source that anchors one of the best-transferring hospital groups.","marker":"Pollard et al. [2018]"},{"why":"Provides HiRID, the high-resolution European ICU source used both as a data contribution and as the fine-tuning study's target.","marker":"Faltys et al. [2021]"},{"why":"Provides PICdb, the pediatric Chinese ICU dataset that supports the paper's claim of incorporating Asian data and pediatric diversity.","marker":"Li et al. [2019]"},{"why":"Provides MIMIC-IV-ED, the only emergency department source, supporting the first ICU+ED harmonized collection.","marker":"Johnson et al. [2023a]"}],"fun_headline_variants":["Harmonized critical care time series: 600k stays, 3 continents","Cross-hospital transfer benchmark for critical care from 600k harmonized stays","Treatment harmonization unlocks cross-hospital ICU time series modeling","600k critical care stays, harmonized treatments, transfer learning ready"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that mapping each hospital's local drug and measurement codes into shared clinical concepts preserves what the treatment actually means across hospitals; if a norepinephrine rate in one database is not the same clinical intervention as in another, the transfer benchmark and the harmonization claim are both confounded.","fun_headline_variants_meta":{"raw":{"variants":["Harmonized critical care time series: 600k stays, 3 continents","Cross-hospital transfer benchmark for critical care from 600k harmonized stays","Treatment harmonization unlocks cross-hospital ICU time series modeling","600k critical care stays, harmonized treatments, transfer learning ready"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000608,"raw_usage":{"total_tokens":2823,"prompt_tokens":926,"completion_tokens":1897,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":1821}},"tokens_in":542,"tokens_out":1897,"duration_ms":12822,"temperature":1.0,"reasoning_tokens":1821,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:12:50.927664+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a core treatment concept such as norepinephrine rate and compare its distribution, its relation to patient physiology, and its association with outcomes across MIMIC-IV, HiRID, SICdb, and UMCdb; if the harmonized rates show systematically different dosing conventions or the same code maps to different actual drugs at different sites, the clinical comparability behind the transfer numbers breaks down.","supporting_citations":[{"cited_title":"Early prediction of circulatory failure in the intensive care unit using machine learning","cited_arxiv_id":null,"evidence_quote":"Defines the circulatory failure early-event prediction task, its 8-hour horizon, and model baselines used in the benchmark."},{"cited_title":"u ser, Xinrui Lyu, Martin Faltys, and Gunnar R\\","cited_arxiv_id":null,"evidence_quote":"Provides HiRID, the high-resolution European ICU source used both as a data contribution and as the fine-tuning study's target."}],"review_version":1}